Specs are the new code review
When agents write the code, humans add value in the spec and the evals. How Even's process runs: scope, build in parallel, review agent, evals, human gate.
Notes from building with agent teams and from working with teams. Short, concrete, opinionated.
When agents write the code, humans add value in the spec and the evals. How Even's process runs: scope, build in parallel, review agent, evals, human gate.
What others do
JPMorgan, Imperial, Klaviyo, Figma, Chinasoft. From a safe chat for employees to agents wired into the systems. What the DACH mid-market should copy and skip.
One day, twelve people, a live demo on real code or a real workflow, a decision at the end. Why that beats a 40-page strategy deck, and how the day runs.
What others do
Spellbook, EvenUp, Kimi, Morgan Stanley. Long-context document work is the quietest big win in enterprise AI. Where it applies outside legal.
What others do
Klarna, Intercom, Autodesk, Starbucks. Read for the mechanism, not the headline: an operator wired into backend APIs, with evals. What a COO should measure.
AI news
Haiku 5.5 moves the small-model price floor, OpenAI's Cookbook shows why tool tests lie, a US open model arrives. What it means for a DACH engineering leader.
Evals, review load and cycle time. If you do not have a baseline for these three before the agents arrive, you will not know whether they helped.
An assistant drafts; a human sends. An operator owns the outcome and escalates when it must. The difference decides whether AI removes work or adds a review step.
Most AI advice comes from people who have not shipped with it. We build Even with agent teams first, then teach. Here is why that order is non-negotiable.
AI news
GPT-6.1 Sol at a fifth of Astra's price, Sonnet 5.5 up to 30% cheaper, Pi 1.0 stable, sandboxes per cloud. What a DACH engineering leader does with it.