Case Studies Best Practices
Ten practices for documenting and learning from your own agent case studies so the next build starts from evidence, not folklore.
Search across all documentation pages
Ten practices for documenting and learning from your own agent case studies so the next build starts from evidence, not folklore.
Use this list when a pilot ends, after a costly incident, or when someone asks "can we do what Team X did?"
done_when, and non-goals before celebrating the demo. Retrofitting a frame invites revisionist success.| Stage | Habits | Exit criterion |
|---|---|---|
| Kickoff | 1-3 | Frame + metrics + architecture sketch exist |
| Pilot end | 4-6 | Evidence links + trade-offs + dated numbers |
| Hardening | 7-8 | Bounds and incidents written |
| Org reuse | 9-10 | Revisit plan + copy gate |
# <Name> - case study (YYYY-MM-DD)
## Problem & done_when
## Architecture (roles)
## Tools & permissions
## Bounds & human gates
## Eval & metrics (window)
## Outcomes (before/after if any)
## Trade-offs rejected
## Incidents / near-misses
## Open questions & revisit date
## Links (ADR, eval, dashboards)Often 60-90 minutes if you captured metrics during the pilot. If it takes days, your instrumentation was missing - fix that next time.
The tech lead of the pilot owns accuracy; a staff+ engineer may own the org template. Ownership beats wiki orphans.
They may produce external versions with legal review. Keep an internal engineering source of truth with full trade-offs and incidents.
Still write it. Failed pilots with clear causes are high-leverage teaching artifacts.
No. Require them for multi-week pilots, production exposure, or patterns others will copy. Spike notes can be ten lines.
ADRs freeze single decisions. Case studies narrate the journey and outcomes. Link them; do not merge them into one muddy doc.
As inspiration only. Their traffic, risk, and ops maturity are not yours. Write the internal version before copying architecture.
Golden-set definition and scores, cost per successful task (or honest proxy), and one redacted trace showing a normal run and a failure run.
At least when model routing, major tools, or success criteria change - and on a calendar cadence (for example quarterly) for production agents.
In-repo docs/case-studies/ or an internal docs space with PR review. Chat threads are not an archive.
Add handoff contracts, per-agent allowlists, and orchestrator budgets to the same skeleton. Do not invent a parallel culture.
Shipping a polished architecture diagram with no metrics, no rejects, and no bounds - then watching another team clone the diagram into an outage.
Stack versions: Pins from the category manifest (verify at build): OpenRouter (~315+ models, July 2026 pricing/fees); LangGraph 1.0+; CrewAI 1.14+; Microsoft Agent Framework 1.0; Vercel AI SDK 6; Pydantic AI (latest); LlamaIndex (latest); OpenAI Agents SDK (latest + MCP); MCP (Linux Foundation governance); A2A (HTTP+SSE+JSON-RPC 2.0); Solana
@solana/web3.js+@solana/spl-token.
Reviewed by Chris St. John·Last updated Jul 16, 2026