Programming
Building with models: the parts that are not the model
A working AI feature is mostly plumbing: context, evaluation, failure handling.
The first thing to internalise: the model is the smallest part of a working AI feature. The interesting engineering is what surrounds it — what context you give it, how you check the output, and what happens when it is wrong.
Context is the product. The same model produces a useless answer and an excellent one depending entirely on what you put in the prompt. For anything beyond a toy, the context is assembled from your data — retrieved, ranked, trimmed to fit — and that assembly logic is where most of the value lives.
A MINIMUM VIABLE AI FEATURE — note what is NOT the model call
1. retrieve pull the few documents/snippets that could matter
(vector search, keyword search, or plain SQL — start simple)
2. assemble order them, trim to a budget, label them clearly
3. call model(prompt, context) <- this line is the small part
4. validate is the output the right SHAPE? parseable? in range?
5. fall back what does the user see if it is wrong or the call fails?
6. measure log inputs/outputs so you can tell if quality moved
If steps 1,2,4,5,6 are missing, you have a demo, not a feature.Evaluation is what separates a demo from a product. Without it, every change is a guess and every regression is discovered by a user. You do not need a research-grade harness to start: a fixed set of 20–50 real inputs with known-good outputs, run automatically on every change, catches most regressions.
What is usually the largest part of a working AI feature?
Context
What you put in the prompt — for real products, assembled from your data.
Evaluation set
A fixed set of real inputs with known-good outputs, run on every change.
Deterministic fallback
What the user gets when the model fails or returns nonsense.
Review cards
Context
What you put in the prompt — for real products, assembled from your data.
Evaluation set
A fixed set of real inputs with known-good outputs, run on every change.
Deterministic fallback
What the user gets when the model fails or returns nonsense.
Sources for this lesson
Below are the references, editions and original links for further reading and checking.
DocsBuilding effective agents / evaluation guidesfree
Anthropic
持续更新
关于「模型之外那些零件」(上下文组装、工具调用、评估、失败兜底)的工程实践指南,比多数教程更贴近真实产品。
BookDeep Learningfree
Ian Goodfellow, Yoshua Bengio, Aaron Courville
MIT Press(免费在线)
深度学习理论基础。第 6 章讲反向传播——本质就是链式法则沿复合链逐层倒着用一遍。
Lights up these nodes in the hub:c-ai