InsightsWhat others do2 min read
Legal and documents: 15 hours to 15 minutes
Spellbook, EvenUp, Kimi, Morgan Stanley. Long-context document work is the quietest big win in enterprise AI. Where it applies outside legal.
Cliff des Ligneris
The least discussed enterprise AI win is also the most repeatable: one person, one long document set, one question.
Second post in the series. Same rules: the figures are the companies’ own public claims, I have not verified them, and wildcard* was not involved in any of these projects.
EvenUp, in Anthropic’s customer stories, says it cut the drafting of a personal injury demand package from 15 hours to 15 minutes. The input is a stack of medical and legal files. The output is a structured document with the evidence pulled out and cited. Nobody chats with it. A paralegal reviews it.
Spellbook, same source, reports 530,000 automated contract reviews a month. The mechanism is long-context analysis of a whole contract against a set of clause expectations, with suggested revisions. Again, no conversation. Document in, marked-up document out.
Kimi Legal & Urban, from Moonshot AI’s enterprise case studies, applies the same idea to Chinese government standards (GB/T) and multi-file reports. The reviewer locates the relevant section across files and inserts precise edits. The interesting part is the multi-file comparison, which is exactly where humans are slowest.
Morgan Stanley is the one that looks like chat but is not. Per the Morgan Stanley press release and OpenAI’s case study, advisors query a library of over 100,000 research reports through a retrieval system on GPT-4. The company claims 98 percent advisor adoption and a jump in document access from 20 to 80 percent. The win is not that advisors can ask questions. It is that the answer comes with the source document attached.
Four different providers, one shape. Long context plus structured output plus a human who checks the citations. No workflow redesign. No change programme. The job stayed the same; the first draft got fast.
That is why this applies far outside legal. Any function where a trained person reads a long document to produce a shorter one is a candidate:
- Procurement: supplier offers against your requirements, deviations listed.
- Compliance and audit: a policy set against a new regulation, gaps marked.
- Tenders: the RFP against your capability list, a first answer per section.
- Technical documentation: a specification against the code’s actual behaviour.
- Insurance and claims: the file against the policy, with the relevant clause quoted.
For a CTO or COO, the test is simple. Pick one document type your team produces at least weekly. Count the hours from first file to first draft. Build the narrowest version: long context, structured output, every claim cited to a page. Measure the hours again, and measure how many citations the reviewer rejected.
If the citation rejection rate is above a few percent, the system is not ready. If it is below, you have EvenUp’s result at your scale, and you built it in weeks.
Where does this break? I want to hear about a document workflow where long context did not help.