How Macrify's discovery engine works ← macrify.me

A mountain of paper. A matter you can question.

Macrify's discovery engine reads an entire discovery production (PDFs, email mailboxes, spreadsheets, scans) and turns it into a fact record you can search, sort, and interrogate. Every answer points to the exact page it came from. You verify by looking, not by trusting.

In development at Macrify · piloted on real litigation Free & open source on release

Discovery isn't a reading problem. It's a finding problem.

A single production can arrive as hundreds of thousands of files: court filings, contracts, email mailboxes with every attachment, spreadsheets, photographs of paper, decade-old scans that no search tool can read. Somewhere in there are the dozen documents that decide the case.

The traditional answer is weeks of attorney and paralegal hours spent opening files one at a time: expensive hours that are mostly looking, not lawyering. The newer answer, handing the pile to an AI and hoping, replaces one problem with a worse one: answers you can't check.

Macrify's engine takes a third path. It does the finding, keeps a receipt for every fact it reports, and leaves the judgment where it belongs: with you.

Every answer shows its receipts.

When the engine tells you something about your case, it doesn't ask to be believed. Each statement carries a citation to a specific document, a specific page, and the specific characters on that page, with the quoted passage stored alongside it. Checking a claim is one click, not an afternoon of re-reading.

Try it. The claims below come from an invented sample dispute, but the mechanics are exactly what the engine does. Click a citation to open its source.

“When did the final inspection actually happen — and who knew?” Worked example · documents invented

The engine's draft answer, as an attorney would see it:

By March 3, the project manager had already moved the final inspection to March 14.

The certificate of completion was signed on March 9 — five days before that inspection.

The site superintendent admitted the inspection was skipped entirely.

Quotes are checked to the character

A stored citation must reproduce the document's text exactly: the same characters, in the same place. Close paraphrases and “that's roughly what it said” do not pass. If the quote doesn't match, the claim is rejected before it's saved.

Citations don't rot

Each citation also keeps the document's digital fingerprint and its exact position, page and character range. If a file were ever swapped or altered, the mismatch surfaces. What you cited is what stays cited.

Independently re-checkable

A single verify command re-reads the whole matter file and re-tests it: file hashes, text, and every citation's anchor. Drift anywhere fails loudly instead of hiding quietly. The record can prove itself again, any day you ask.

Two honesty rules run through everything the engine produces. Relevance scores are ranking signals, not conclusions: they order the reading, they never decide the case. And every generated work product is stamped as a draft for attorney review. The engine finds and cites; lawyers judge.

Explore the matter from any angle.

Once the corpus is read, the facts don't sit in a chat window. They become durable work products you can open, print, and hand to co-counsel. Each one is a single self-contained file, every entry pinned to its source passage, with a note stating exactly what was included and what was left out.

The master chronology

Every dated fact in the record, in order, each with its citation.

2022-11-18Phase II contract executed; substantial-completion date set. RVB-0000214 · p.2
2023-02-24Storm damage noted on punch-list walkthrough. RVB-0003967 · p.1
2023-03-03Final inspection moved to March 14. RVB-0004182 · p.3
2023-03-09Certificate of completion signed. RVB-0009117 · p.1
2023-03-14County inspection recorded; three items flagged. RVB-0009544 · p.2
2023-04-02“Final” budget circulated with prior contingency figure. RVB-0007731 · p.1
2023-05-16First warranty claim letter received from the owner. RVB-0010208 · p.1
2023-06-01Demand to repair; dispute counsel copied. RVB-0010461 · p.2

A timeline for each person

What one custodian said, sent, and signed, in sequence.

Mar 03K. Otero receives the reschedule email. RVB-0004182
Mar 08Forwards certificate draft “for signature tomorrow.” RVB-0008990
Mar 09Signs the certificate of completion. RVB-0009117

The hot-doc queue

Documents ranked against your objective, with documents that cut against you in their own lane.

Ranked for review
Reschedule email RVB-0004182
Certificate of completion RVB-0009117
Adverse — flagged separately
Owner's delay-waiver letter RVB-000602

The claim chart

Each element of each claim, mapped to the evidence that speaks to it.

Breach — Element 2: Failure to perform

Certificate signed before the rescheduled inspection. RVB-0009117 · p.1 RVB-0004182 · p.3

Breach — Element 3: Notice

Punch-list walkthrough memo delivered to all parties. RVB-0003967 · p.1

Damages — Element 1: Causation

Flagged inspection items match warranty claims. RVB-0009544 · p.2

Defense — Waiver (adverse)

Owner's letter arguably excusing the delay. RVB-0006020 · p.1

And the matter stays queryable.

Underneath the work products is an indexed record of the whole corpus: full-text search across every document, the people and organizations in it, the dates and events they connect to. “What did we tell the owner's rep in March?” is a lookup with citations, not a reading assignment.

You can also point your own AI assistant at the matter: read-only, with fixed permissions the operator sets up front. It answers under the same citation rules, every disclosure is logged, and documents flagged as potentially privileged are automatically withheld from any cloud-facing view.

Views that hold up later.

Every export records the exact selection that produced it (the filters, the date ranges, the thresholds) and lists its omissions visibly instead of hiding them.

Open the chronology in six months, or hand it to opposing counsel's-eyes-only review, and it still says precisely what it covered, what it excluded, and where every line came from.

And when it's your turn to produce.

Discovery runs both ways. When you owe the other side documents, the engine assembles the outbound set from decisions already on the record: the documents you've tagged to your issues and coded responsive, with email families kept together, and anything marked privileged pulled out automatically — its privilege-log entry drafted from the record, not retyped from memory.

The engine's suggestions stay suggestions. It can rank a candidate against your objective and attach the exact quoted passage that earned the rank; it cannot code a document. Responsiveness and privilege calls are entered by a named reviewer, one document at a time, on an append-only record that keeps every decision's author and history.

And the last mile is deliberately out of the software's hands. Privilege, responsiveness, redaction approval, blocker waivers, final production approval, and delivery are attorney decisions the engine cannot make — or waive. An unresolved privilege flag doesn't get a best guess; it freezes the production until a person rules, and even a deliberate override requires a named reviewer and a written reason, kept on the record. The engine stamps the Bates numbers, builds the load file, and verifies every byte of the finished set. A lawyer decides that it ships.

Outbound production · PROD-002STAGED — NOT APPROVED
RVB-0000214 Phase II contractBates reserved: PRD-000001 – 000012 responsive · attorney-coded
RVB-0004182 Reschedule email, D. Marsh to K. Oteroengine: ranked relevant to the objective · quote attached proposed · review pending
RVB-0008990 Certificate draft + attachmentkept with its parent email family · produced together
RVB-0006294 Email thread with outside counselprivilege-log entry drafted for attorney confirmation withheld · privilege

Six decisions never leave the attorney. The engine stages, numbers, and verifies this set; it cannot approve it, waive a blocker on it, or send it. privilege · responsiveness · redactions · blocker waiver · final approval · delivery

A staged set from the invented Riverbend example. The numbers are illustrative; the mechanics — staged designations, automatic privilege withholding, frozen-until-approved — are the engine's.

You control the AI spend.

AI reading costs money, and on a big corpus the meter matters. The engine treats model usage the way a firm treats any other disbursement: estimated in advance, capped in writing, spent only where the case needs it, and itemized afterward.

The estimate comes first.

Before any model reads anything, the engine profiles the corpus and produces a processing plan: what's in the pile, how it will be handled, and a projected cost range for each AI step, stated as a range rather than a bill.

The unglamorous work never hits the meter at all: making scans searchable, cataloguing files, deduplicating, and filing families all run locally on your own machine, at no per-page charge to anyone.

Illustrative plan: the counts and dollars are invented for this example. Real plans project ranges from your corpus and your chosen models' prices.

A hard cap, in writing.

Set a dollar budget on any run. The engine reserves the projected cost of each batch before sending it; at the cap, the default is to pause and wait for you: finished work is kept, nothing further is spent, and the run resumes when you say so.

Every token and dollar is written to the run's ledger as it happens, so the AI bill is itemized like any other cost of the matter.

Aim it with a sentence.

You give the matter a plain-language objective. The engine ranks every document against it, so expensive attention lands on what your case turns on, instead of paying to read everything with equal care. Change the objective, and the ranking follows.

It's also model-agnostic: pick the model for each job (a fast one for bulk reading, a careful one for ranking), swap by provider or price, or use a subscription seat you already pay for. No model is load-bearing.

Matter objective, written by the attorney
“Surface anything bearing on when the inspection happened, who scheduled it, and who knew the certificate predated it.”
readingfast bulk modelswappable
rankingcareful modelswappable
OCR & filingyour own machine$0

Small jobs, checked work.

Imagine one brilliant clerk handed a warehouse of boxes and told: read everything, remember everything, then tell me what matters. However brilliant, the answer arrives unverifiable: you can't tell the remembered from the imagined. That is, roughly, what “paste the documents into a chatbot” does, at scale, with your case.

This engine is built on the opposite model: a room full of clerks, each given one small, bounded task: read this one document, pull the dates from this one page, compare these two names. Each comes with defined inputs, defined outputs, and a supervisor's ledger recording every completion. The AI is used the same way a good clerk is: for one task at a time, with its work checked before it's filed.

The pipeline's spine (cataloguing, custody, text, filing) is deterministic: run it twice, get the same answer twice. The AI sits inside that spine as a replaceable part, never as the foundation. Each document's progress is written to the ledger stage by stage, so a run interrupted overnight resumes exactly where it stopped, and only the piece that failed is retried, not the whole matter.

This discipline is why the two things lawyers care about move together. Accuracy: small steps produce checkable claims, and every claim is verified against the source before it's stored. Cost: nothing is re-read, models see only what their one task requires, and there is no monolithic read-the-whole-corpus prompt anywhere in the system.

Profile what's in the pile Plan cost, in advance Ingest custody & families Make searchable local OCR, $0 Read & extract one doc per task Merge one master record Work products timelines, hot docs

The work is split into batches of a few thousand documents, processed independently, then merged deterministically into one master record. Big matters don't need a bigger prompt, just more small, checked batches.

The supervisor's ledger

Every document's stage is on the record, and stays there. Stop the run at midnight, start it Tuesday: nothing is repeated, nothing is lost, and the one file that failed is retried on its own.

Why the discipline pays

  • Interruptions are routine, not disasters. The ledger makes every run resumable across nights, weekends, and hardware.
  • Problems surface; they don't disappear. An unreadable file or a locked mailbox becomes a logged exception that a named person must rule on. The software cannot waive it, by design.
  • The record is append-only. Custody events, decisions, and dispositions are written once and kept, in tamper-evident chains: a paper trail you can stand behind.
  • Everything re-proves itself. One command re-checks the whole matter (hashes, text, citations) any time you need to show your work.