Evaluating LLM features without fooling yourself
A practical evaluation stack: golden sets, LLM-as-judge with its known biases, regression gates in CI, and the metrics that actually correlate with user trust.
10 min readInsights
Written by the people doing the engineering, aimed at the person who has to make the call.
A practical evaluation stack: golden sets, LLM-as-judge with its known biases, regression gates in CI, and the metrics that actually correlate with user trust.
10 min readToken pricing is the smallest line in the budget. Here is the cost model we use with clients, covering inference, retrieval infrastructure, evaluation, human review and the ongoing maintenance nobody scopes.
7 min readLet's build together
A short conversation is usually enough to tell whether there is a real case here, and we will say so if there is not.