How we think about production AI.
Engineering deep-dives from the team — on retrieval, grounding, agentic systems, evaluation, and the unglamorous work of getting AI past a security review and into production.
Engineering deep-dives from the team — on retrieval, grounding, agentic systems, evaluation, and the unglamorous work of getting AI past a security review and into production.
How to give AI agents real tool access safely — least-privilege scoping, the Model Context Protocol as a boundary, human-in-the-loop approvals, and auditing every call.
RetrievalNaive retrieval-augmented generation still hallucinates. The four retrieval and grounding techniques — hybrid search, re-ranking, GraphRAG and grounding guardrails — that make enterprise RAG accurate and cited.
EvaluationWhy traditional testing fails for language models, and how to build an evaluation harness that scores grounding, accuracy and safety on every release — not just at launch.
New deep-dives roughly every few weeks. Want one on a specific topic?
Suggest a topic