AI & ML4 min
Evaluating LLM features before you launch them
Shipping an AI feature without an evaluation harness is shipping a system nobody can prove works. Here is the eval setup we build on every engagement.
Insights
Architecture, performance, and how we actually ship — written by the people on the project.
Shipping an AI feature without an evaluation harness is shipping a system nobody can prove works. Here is the eval setup we build on every engagement.
Big-bang rewrites fail for structural reasons, not technical ones. A field guide to incremental extraction from systems that cannot go offline.
We instrumented six e-commerce clients to correlate LCP and INP with conversion. The numbers were larger than any of them expected.

30 minutes · CET · no deck
Bring the paper, the WhatsApp thread, or the spreadsheet. We will tell you what to build first — and what not to.