Skip to content
AI Inside the WorkflowShippedAI

Lab reports to named markers, transcribed and never interpreted

A model reads the lab report as data entry, not analysis. Values stay strings, only the lab's own flags count, and the schema has nowhere to put a diagnosis.

  • Supabase
  • Custom

The problem

A lab report is the one input where a wrong digit or an invented flag flows silently into the plan and nobody notices until a client does. The default behavior of a capable model is the opposite of what you want here: it summarizes, it interprets, it helpfully flags what it thinks is high.

What we built

The model runs on the largest tier by choice. Volume is a handful of documents per intake, and a misread digit costs more than the call. The prompt frames the job as data entry: transcribe only, never infer a flag, never invent a marker, omit rather than guess, and leave out names and dates of birth.

The schema does the enforcing. Each marker has a name, a value stored as a string so results like "less than 0.1" or "1:40" survive, a unit, the reference range as printed, the panel, and a flag from a closed list of five or nothing at all. The collection date is the one on the report, not the upload date. A test asserts the schema has no field that could carry an interpretation, and the plan stages carry a matching rule: labs are described observationally, only the lab's own flags count, and a flag belongs only to the marker that carries it.

The parser is tolerant about packaging and strict about content: a payload that fails the schema is a failed extraction staff can retry, never a half-filled panel, and a model refusal is recorded as its own status. Extraction runs in the background after upload, the Generate button waits for it, and a failed document does not block the plan. The lab summary the plan stages read is built once, deterministically, and reused byte for byte on every regeneration.

Where AI does the work

A model on the largest tier reads each uploaded lab report as data entry and returns named markers with the value, unit, reference range, panel, and collection date exactly as printed, plus the lab's own flag when one is printed. It is told to omit rather than guess.

Where rules do the work

A schema with nowhere to put an interpretation: values are strings so results like less-than signs and ratios survive, the flag is a closed list, and a test asserts no assessment or diagnosis field exists. A payload that fails the schema is a failed extraction staff can retry, never a partial panel. File type is decided by the file's bytes, not the type the upload declares. The summary the plan stages read is built once, deterministically, and reused byte for byte on regeneration.

Outcome

One real five-page panel diffed marker by marker against the PDF text: 32 of 32 markers transcribed, 0 invented, 3 flags matching the report's own 3, collection date correct, no client identifiers. That run caught a flag attributed to two markers when only one carried it; the rule was tightened and a regression test added. One panel is a check, not a benchmark.

Who this fits

  • Invoices, contracts, and forms where a misread digit propagates silently
  • Teams that need a model to transcribe a document without offering an opinion on it
  • Anyone who has been told extraction needs an OCR vendor first

See it running first.

Proof first. Then we'll talk about your stack.

How we work →