Lawsuits, audits, investigations and journalists with a leaked archive all face this: a pile too big to read and a request about what’s in it. The request is only words. What counts is how the lawyer in charge reads them, and hunch is for turning that reading into a question you can run, measure and change.
Try it first
In 2010 a senior lawyer played this role for a public exercise, the TREC Legal Track, and wrote down what the drilling request covers. Would the lawyer count these? The words say drilling. The lawyer meant the whole business of getting oil and gas out of the ground, moved and paid for, as the company did it, and not the talk around it.1. Get the emails
weights: tells hunch how many each one speaks for, so every number on this page is about the whole collection.
2. Ask the request, word for word
prototype/examples/discovery/request_only/drilling.yml
3. Write down what the lawyer meant
Add the lawyer’s reading to the question, in a paragraph (condensed from eight pages of guidance):prototype/examples/discovery/drilling.yml
hunch diff asks it of every email and compares:
4. Decide how much a person reads
Nobody hands over what a model picked without reading it. The model’s job is the other side: setting aside the emails it is sure are not wanted, so a person reads the rest.act says how sure a “no” must be to be set aside, and min_recall says how much of what the lawyer wanted must still reach a person:
prototype/examples/discovery/privileged.yml
hunch test checks the worst plausible case, the bottom of the 95% interval, not the average:
The lesson
The teams in the 2010 exercise, e-discovery firms and universities with up to ten hours of the same lawyer’s time, hit the same wall on drilling. None found more than 25% of what the lawyer wanted, and the organisers wrote that none captured the lawyer’s broad reading of the request. A request is not a question until someone writes down what it means. In hunch that writing is the spec: it is read, diffed and tested like code, so when the lawyer’s reading changes you see which documents move before anyone relies on them.Use it on your data
Pointsource: at your documents, put the request and your review protocol in instructions, have a lawyer judge a few hundred, and hunch test tells you how much a person must read.
How it was measured
The data
The data
The TREC 2010 Legal Track Interactive task: four requests for production from a mock complaint (drilling, responses to spills, lobbying, and a privilege review), over the EDRM Enron Email Data Set v2 (ZL Technologies, CC BY 3.0 US), 455,449 messages. Professional review firms judged a stratified sample of each topic; the topic’s senior lawyer (the “Topic Authority”) then adjudicated about 10% of those calls, the ones teams appealed plus a sample of the rest. This page uses the final judgments, 25,353 in all, leaving out 154 marked unreadable. Each judgment carries the probability its email was sampled; summing one over it gives the collection’s 455,449 for every topic (within 0.3% once the unreadable ones are left out), and
weights: in each spec reweights by it. fetch.py keeps the first 20,000 characters of an email and 10,000 of its attachments; the specs send 6,000 and 2,000. The emails name real people; they are public records, and this page quotes no email and names no one.The results, in full
The results, in full
From
measure.py, weighted to the collection, 95% intervals from 1,000 bootstrap resamples within strata; hunch says yes at p(yes) ≥ 0.5.Keywords, as regular expressions on the lowercased email and attachments: drilling
drill|\boil\b|\bgas\b|extraction|pipeline|reserves|exploration; privilege privilege|attorney|counsel|lawyer|legal|litigation|lawsuit|confidential; lobbying lobby|legislat|senator|congress|regulat|governor|\bbill\b|testimony|government affairs. They read 17%, 14% and 13% of the collection; hunch’s ranking, reading as much, finds 84%, 84% and 94%. (A first, narrower query, drill|oil and gas|\bwell\b|\brig\b, found 19% on drilling: a keyword search is only as good as its words.) On lobbying the request’s words were already enough; the lawyer’s reading moved precision up and recall down. To find 80% by ranking alone, a person reads the top 12.7% (9.3–17.4) for drilling, 12.2% (10.4–14.7) for privilege, 4.5% (3.8–5.1) for lobbying. The min_recall checks use a Wilson interval on the effective sample size (941 privileged emails weigh as 453).The recall bars were chosen on the same answers
The recall bars were chosen on the same answers
Each spec’s
act for “no” (0.82 drilling, 0.75 privilege, 0.65 lobbying) is the lowest that hunch test reported as keeping recall at 80% on the lower bound, found on the same judged emails it is then checked against, so the PASS is optimistic. A fair check: choose the bar on half the emails (split by a hash of the id) and test it on the other half. Chosen on one half: 0.86, 0.77, 0.65; on the other half recall is 89.9% (83.3–94.1), 87.9% (83.2–91.4) and 83.3% (78.5–87.2). Drilling and privilege hold; lobbying’s lower bound falls just under 80%. On your own documents, choose the bar on one set of reviewed emails and check it on another.Why spills is left out
Why spills is left out
Only 184 of the judged emails were about responses to spills, and three of them, sampled from the part of the collection no team flagged, stand for most of the 575 spill emails the sample implies are in the collection. Weighted, the 184 count as 8. hunch’s recall there is somewhere between 9% and 44%, and
hunch test says no recall bar can be promised; the spec is in the example without one.The human reviewers
The human reviewers
The first-pass reviewers’ calls are published too, and against the final judgments they score far higher (F1 83% on drilling, 87% on privilege). That comparison is largely circular: the final judgments are the reviewers’ calls, and only about 10% of them went to the senior lawyer (appeals, plus 884 others the organisers chose, of which about a quarter were overturned). Everywhere else the reviewer’s call is the answer key. It is an upper bound on the reviewers, not a measurement. For a fair comparison of people and software on the 2009 edition of this exercise, see Grossman and Cormack, “Technology-Assisted Review in E-Discovery Can Be More Effective and More Efficient Than Exhaustive Manual Review”, Richmond Journal of Law and Technology, 2011.
Against the 2010 teams, and why it isn't a race
Against the 2010 teams, and why it isn't a race
The track’s overview reports each team’s final result: the best F1 was 26% on drilling (no team above 25% recall), 67% on lobbying and 41% on privilege, against hunch’s 53%, 64% and 48% above. It is not a fair race. The paragraph each spec adds is our condensation of the instructions the reviewers who produced the judgments worked from, compiled largely from what each lawyer told the teams (the published version is dated after the teams had submitted). That is the answer key’s own instrument, a real advantage the teams didn’t have in that form; they questioned the lawyer directly, some for up to ten hours, some not at all. The lobbying paragraph also keeps two names specific to this case from those instructions (the Independent Energy Producers Association and the CPUC). A 2026 model is being compared with 2010 systems, and the Enron emails are all over the web (the judgments much less so). What the comparison does show is the cost: about $0.33 and under two minutes per topic.
Cost
Cost
Both wordings of all four topics, every judged email: 0.