Working notes · Azure & AWS

Grounding, agents and evaluation - with the numbers attached.

Notes from production deployments on Copilot Studio, Azure AI Search, Azure AI Foundry and Amazon Bedrock. Retrieval quality, agent routing, eval gates, and the invoices that follow - measured before claimed.

15
Posts
4
Platforms
2025
Since

All writing

15 posts · newest first
09 JulCopilot Studio agents that actually ground: a knowledge-source auditSix sources, one question set, and the three configurations that produced answers nobody had to double-check.Copilot Studio·Field note·Retrieval, Governance1 min24 JunFoundry evaluations in CI: gating a prompt change like codeA pull request that changes a system prompt should fail the same way a bad migration does.Azure AI Foundry·Tutorial·Evaluation1 min05 JunBedrock knowledge bases versus rolling your own retrievalWhere the managed pipeline is a bargain, where it quietly caps your recall, and what the migration costs.Amazon Bedrock·Deep dive·Retrieval, Cost1 min21 MayBedrock guardrails, measured: what they catch and what they costTwo thousand adversarial prompts, one policy set, and the latency you pay per intervention.Amazon Bedrock·Case study·Governance, Evaluation1 min06 MayTopic triggers in Copilot Studio: the routing layer nobody documentsWhy your agent answers the wrong question, and how to make routing observable before users find it.Copilot Studio·Field note·Agents1 min19 AprHybrid retrieval and the semantic ranker: when it stops payingThe ranker earns its price on ambiguous questions and almost nothing on lookups. Here is the crossover.Azure AI Search·Deep dive·Retrieval, Cost1 min02 AprCite or it did not happen: enforcing citations in tool-calling agentsA grounding contract, a validator, and a refusal path - an agent that cannot cite should decline.Azure AI Foundry·Essay·Agents, Governance1 min17 MarPrompt flow to Foundry agents: a migration in four commitsWhat mapped cleanly, what had to be rebuilt, and the eval set that made the swap boring.Azure AI Foundry·Case study·Agents, Evaluation1 min26 FebCross-cloud RAG: one corpus, two retrievers, honest numbersThe same documents indexed on Azure AI Search and Bedrock, scored on the same questions.Amazon Bedrock·Deep dive·Retrieval, Evaluation1 min04 FebVector index size is a bill, not a metricDimensions, quantisation and replica count - where the money actually goes in a production index.Azure AI Search·Field note·Cost1 min15 JanThe 200-question eval set that replaced our vibe checksHow to build a labelled set in a week, and why 200 good questions beat 5 000 synthetic ones.Azure AI Foundry·Tutorial·Evaluation1 min09 DecWhat Copilot Studio taught me about scoping an assistantEvery agent that failed review failed for the same reason: it was asked to be useful to everyone.Copilot Studio·Essay·Agents, Governance1 min22 OctSemantic ranker: a before-and-after on 3 400 queriesThe first honest measurement I made on this stack, and the one that started the eval set.Azure AI Search·Case study·Retrieval, Evaluation1 min03 SepOur first RAG pipeline was a search problem in disguiseWe shipped an assistant, then discovered we had shipped a bad search engine with a friendly voice.Amazon Bedrock·Essay·Retrieval1 min
Full archive →
ESC
Move OpenT Theme