> ## Documentation Index
> Fetch the complete documentation index at: https://fuguai.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Find what a lawsuit asks for

> 455,449 Enron emails, requests from a real legal exercise, and a question that reads them the way the senior lawyer does.

export const DISCOVERY_READ = {
  "drilling": {
    "title": "Oil and gas drilling",
    "total": 455404,
    "relevant": 18973,
    "protocol": [0.0, 9.5, 19.2, 25.1, 31.2, 35.8, 42.3, 47.6, 51.6, 54.6, 59.2, 62.5, 65.0, 66.5, 69.1, 71.8, 72.8, 73.5, 75.0, 75.3, 75.5, 76.4, 77.7, 78.0, 79.7, 79.8, 80.3, 81.3, 81.3, 82.4, 83.1, 83.2, 83.4, 84.3, 84.3, 84.5, 85.3, 85.3, 86.1, 86.4, 86.4, 86.7, 86.8, 87.0, 87.3, 87.3, 87.5, 87.5, 88.3, 88.3, 88.3, 88.4, 89.3, 89.4, 89.4, 90.3, 90.5, 90.5, 90.5, 90.6, 90.7, 91.4, 91.5, 91.5, 91.7, 91.7, 91.8, 91.8, 92.8, 92.8, 92.9, 93.0, 93.0, 93.1, 93.3, 93.3, 93.4, 93.5, 93.5, 93.6, 93.6, 93.6, 93.6, 93.6, 94.6, 94.6, 94.6, 94.6, 94.6, 94.6, 94.6, 94.7, 94.7, 94.8, 95.8, 95.8, 95.8, 97.3, 97.3, 97.3, 97.4, 97.4, 97.6, 97.6, 97.6, 97.6, 97.6, 97.6, 97.7, 97.7, 97.8, 97.8, 97.9, 98.1, 98.1, 98.2, 98.3, 98.3, 98.3, 98.3, 98.3, 98.3, 98.4, 98.4, 98.5, 99.0, 99.1, 99.1, 99.1, 99.1, 99.1, 99.1, 99.1, 99.1, 99.1, 99.1, 99.1, 99.1, 99.1, 99.1, 99.1, 99.2, 99.2, 99.2, 99.2, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0],
    "request": [0.0, 5.9, 7.9, 9.2, 10.5, 11.1, 12.3, 13.3, 14.4, 16.0, 17.7, 18.0, 18.5, 20.4, 20.6, 22.3, 22.6, 23.9, 24.1, 24.8, 25.8, 26.0, 29.3, 30.2, 31.5, 31.9, 33.4, 34.4, 35.8, 36.8, 36.9, 36.9, 38.1, 39.3, 39.6, 39.7, 40.0, 41.7, 42.4, 42.4, 43.2, 43.8, 44.6, 46.2, 47.4, 50.0, 50.2, 50.6, 51.1, 51.7, 52.1, 53.1, 55.0, 55.9, 57.9, 58.7, 59.2, 59.5, 61.7, 62.1, 62.6, 63.1, 64.0, 64.2, 64.9, 64.9, 66.3, 67.1, 67.1, 67.5, 67.7, 68.9, 69.4, 70.5, 70.6, 70.9, 71.0, 71.6, 73.2, 73.3, 74.3, 74.6, 74.7, 75.3, 76.3, 76.4, 76.7, 76.8, 76.8, 77.3, 78.0, 78.1, 79.7, 80.6, 80.9, 81.7, 81.8, 82.2, 82.3, 84.7, 85.0, 85.3, 85.4, 86.2, 87.7, 87.7, 87.8, 88.0, 88.8, 88.9, 89.0, 90.0, 90.9, 91.6, 91.7, 91.7, 91.9, 91.9, 91.9, 92.0, 92.1, 93.2, 93.3, 94.0, 94.2, 94.2, 94.2, 96.1, 97.1, 97.1, 97.1, 97.1, 97.3, 97.4, 98.1, 98.1, 98.1, 98.2, 98.2, 98.2, 98.2, 98.2, 98.2, 98.2, 98.2, 98.3, 98.3, 98.3, 98.4, 98.4, 98.4, 98.4, 98.4, 98.5, 98.5, 98.5, 99.5, 99.8, 99.8, 99.8, 99.8, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0],
    "keywords": [17.0, 91.0]
  },
  "privileged": {
    "title": "Privileged",
    "total": 455237,
    "relevant": 20176,
    "protocol": [0.0, 9.0, 17.1, 23.0, 28.4, 32.9, 38.2, 41.7, 46.7, 50.7, 53.4, 56.7, 59.5, 62.4, 65.7, 67.6, 68.9, 71.4, 72.8, 73.7, 74.5, 75.7, 77.9, 78.4, 79.5, 81.0, 81.8, 82.6, 83.6, 83.6, 84.2, 85.0, 85.7, 85.9, 87.2, 87.2, 87.4, 87.7, 88.1, 88.2, 88.5, 88.7, 89.4, 89.6, 89.8, 90.6, 92.1, 92.9, 93.0, 93.0, 93.0, 93.0, 93.1, 93.1, 93.1, 93.4, 93.5, 93.6, 93.7, 93.7, 94.5, 95.1, 95.3, 95.3, 96.1, 96.1, 96.1, 96.1, 96.3, 96.3, 96.4, 96.4, 96.5, 96.6, 96.7, 96.7, 96.7, 96.7, 96.8, 96.8, 96.8, 96.8, 96.8, 96.8, 96.8, 96.8, 96.8, 97.0, 97.0, 97.1, 97.1, 97.1, 97.1, 97.1, 97.4, 97.4, 97.5, 97.5, 97.5, 97.5, 97.5, 97.5, 97.5, 97.5, 97.5, 97.5, 97.5, 97.5, 98.2, 98.2, 98.2, 98.2, 98.3, 98.3, 98.3, 98.3, 98.3, 98.3, 98.3, 98.9, 98.9, 98.9, 98.9, 99.0, 99.0, 99.0, 99.1, 99.1, 99.1, 99.1, 99.1, 99.1, 99.2, 99.8, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0],
    "request": [0.0, 8.7, 15.9, 20.6, 24.0, 26.9, 30.9, 34.8, 37.8, 41.1, 44.7, 48.4, 51.3, 52.8, 55.3, 56.0, 58.3, 61.0, 61.9, 63.5, 64.4, 65.7, 66.7, 67.3, 68.3, 68.6, 70.5, 70.9, 71.6, 73.3, 74.4, 74.4, 75.7, 77.0, 77.8, 77.8, 78.1, 78.8, 79.5, 80.1, 80.2, 80.5, 81.3, 81.7, 82.6, 82.7, 82.9, 83.0, 83.9, 84.2, 84.6, 84.8, 85.0, 85.0, 85.5, 85.9, 86.0, 87.3, 87.3, 87.4, 87.6, 87.7, 87.7, 87.9, 88.7, 88.9, 89.0, 89.0, 89.1, 89.3, 89.3, 89.3, 89.4, 89.8, 89.9, 90.0, 90.2, 90.2, 90.3, 90.9, 90.9, 91.2, 91.7, 91.9, 92.0, 92.4, 92.5, 93.4, 94.0, 94.0, 94.0, 94.1, 94.4, 94.4, 94.4, 95.2, 95.2, 95.5, 95.5, 95.5, 95.6, 95.8, 95.8, 95.9, 96.0, 96.1, 96.3, 96.3, 96.3, 96.9, 96.9, 96.9, 97.1, 97.1, 97.1, 97.1, 97.1, 97.1, 97.3, 97.4, 97.4, 97.5, 97.5, 97.5, 97.5, 98.1, 98.1, 98.2, 98.2, 98.4, 98.4, 98.4, 98.4, 98.4, 98.5, 99.1, 99.1, 99.1, 99.2, 99.2, 99.2, 99.2, 99.2, 99.2, 99.2, 99.2, 99.2, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.3, 99.4, 99.4, 99.4, 99.4, 99.4, 99.4, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0],
    "keywords": [14.2, 66.8]
  },
  "lobbying": {
    "title": "Lobbying",
    "total": 454352,
    "relevant": 12124,
    "protocol": [0.0, 18.2, 33.3, 45.7, 55.5, 60.1, 68.2, 73.2, 77.4, 80.0, 84.1, 85.7, 87.8, 88.6, 88.9, 90.6, 91.0, 91.4, 91.7, 92.3, 92.5, 92.6, 92.9, 93.7, 93.8, 94.3, 94.3, 94.5, 94.9, 95.2, 95.6, 95.6, 95.8, 96.0, 96.1, 96.2, 96.4, 96.4, 96.4, 96.4, 96.4, 96.4, 96.5, 96.5, 96.8, 96.9, 96.9, 97.0, 98.3, 98.3, 98.3, 98.3, 98.6, 98.6, 98.6, 98.6, 98.7, 98.8, 98.9, 98.9, 98.9, 98.9, 98.9, 99.1, 99.1, 99.1, 99.2, 99.2, 99.3, 99.3, 99.3, 99.3, 99.4, 99.4, 99.5, 99.5, 99.5, 99.5, 99.5, 99.5, 99.5, 99.5, 99.5, 99.6, 99.6, 99.6, 99.6, 99.6, 99.6, 99.6, 99.7, 99.7, 99.7, 99.7, 99.7, 99.7, 99.7, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0],
    "request": [0.0, 18.2, 32.8, 45.2, 54.4, 62.8, 69.9, 72.8, 79.3, 82.8, 85.9, 87.9, 89.1, 90.0, 91.9, 93.0, 93.4, 93.9, 94.7, 94.7, 95.2, 95.8, 96.2, 96.2, 96.3, 96.4, 96.4, 96.4, 97.8, 97.8, 98.2, 98.3, 98.3, 98.3, 98.5, 98.5, 98.5, 98.5, 98.6, 98.6, 98.6, 98.6, 98.6, 98.7, 98.7, 98.7, 98.7, 98.7, 98.8, 98.8, 98.8, 98.9, 99.0, 99.1, 99.1, 99.1, 99.1, 99.3, 99.3, 99.3, 99.3, 99.4, 99.4, 99.4, 99.4, 99.5, 99.5, 99.5, 99.5, 99.6, 99.6, 99.6, 99.6, 99.6, 99.6, 99.6, 99.6, 99.6, 99.6, 99.6, 99.6, 99.6, 99.6, 99.7, 99.7, 99.7, 99.7, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.8, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 99.9, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0, 100.0],
    "keywords": [13.1, 77.3]
  }
};

export const DISCOVERY_FOUND = [["The request's words", 7.9], ["The lawyer's reading", 63.3]];

export const ScopePuzzle = ({items}) => {
  const [picks, setPicks] = useState({});
  const teal = "#7D969B", red = "#C64D35", grey = "rgba(128,128,128,0.25)";
  const done = Object.keys(picks).length === items.length;
  const agree = items.filter((it, k) => picks[k] === it.yes).length;
  const btn = (k, v, label) => <button key={label} onClick={() => picks[k] === undefined && setPicks({
    ...picks,
    [k]: v
  })} disabled={picks[k] !== undefined} style={{
    border: `1px solid ${picks[k] === v ? "currentColor" : grey}`,
    borderRadius: 8,
    padding: "3px 12px",
    background: "transparent",
    color: "inherit",
    font: "inherit",
    fontSize: 13,
    cursor: picks[k] === undefined ? "pointer" : "default"
  }}>{label}</button>;
  return <div className="not-prose hunch-widget" style={{
    border: `1px solid ${grey}`,
    borderRadius: 12,
    padding: 16,
    margin: "16px 0"
  }}>
      <div style={{
    fontSize: 13,
    opacity: 0.7
  }}>Does the lawyer count it as relating to oil and gas drilling?</div>
      {items.map((it, k) => {
    const shown = picks[k] !== undefined;
    return <div key={k} style={{
      padding: "10px 0",
      borderTop: k ? `1px solid ${grey}` : "none"
    }}>
            <div style={{
      display: "flex",
      justifyContent: "space-between",
      alignItems: "center",
      gap: 10,
      flexWrap: "wrap"
    }}>
              <span style={{
      fontSize: 15,
      flex: "1 1 240px"
    }}>{it.doc}</span>
              <span style={{
      display: "flex",
      gap: 6
    }}>{btn(k, true, "Yes")}{btn(k, false, "No")}</span>
            </div>
            {shown && <div style={{
      fontSize: 13,
      marginTop: 6,
      lineHeight: 1.5
    }}>
                <b style={{
      color: picks[k] === it.yes ? teal : red
    }}>{it.yes ? "Yes." : "No."}</b> {it.why}
              </div>}
          </div>;
  })}
      {done && <p style={{
    fontSize: 14,
    margin: "10px 0 0"
  }}>
          You read it as the lawyer did on <b>{agree} of {items.length}</b>. The request's words don't say which way any of these go.
        </p>}
    </div>;
};

export const ReadSlider = ({data}) => {
  const names = Object.keys(data);
  const [topic, setTopic] = useState(names[0]);
  const [step, setStep] = useState(24);
  const d = data[topic];
  const teal = "#7D969B", red = "#C64D35", gold = "#C9A227", grey = "rgba(128,128,128,0.25)";
  const W = 320, H = 170, L = 8, B = 8, maxStep = 100;
  const x = i => L + i / maxStep * (W - L - 4), yv = v => H - B - v / 100 * (H - B - 6);
  const path = arr => arr.slice(0, maxStep + 1).map((v, i) => `${i ? "L" : "M"}${x(i).toFixed(1)},${yv(v).toFixed(1)}`).join("");
  const kx = Math.min(d.keywords[0] * 2, maxStep);
  const read = step / 2, people = Math.round(d.total * read / 100 / 1000) * 1000;
  return <div className="not-prose hunch-widget" style={{
    border: `1px solid ${grey}`,
    borderRadius: 12,
    padding: 16,
    margin: "16px 0"
  }}>
      <div style={{
    display: "flex",
    gap: 6,
    flexWrap: "wrap",
    marginBottom: 10
  }}>
        {names.map(n => <button key={n} onClick={() => setTopic(n)} style={{
    border: `1px solid ${n === topic ? "currentColor" : grey}`,
    borderRadius: 8,
    padding: "4px 10px",
    background: "transparent",
    color: "inherit",
    font: "inherit",
    fontSize: 13,
    cursor: "pointer",
    opacity: n === topic ? 1 : 0.75
  }}>{data[n].title}</button>)}
      </div>
      <label style={{
    display: "block",
    fontSize: 14
  }}>
        A person reads the top <b>{read}%</b> of the collection, about {people.toLocaleString("en")} messages
        <input type="range" min={0} max={maxStep} value={step} onChange={e => setStep(+e.target.value)} aria-label="Share of the collection a person reads" style={{
    width: "100%",
    accentColor: teal,
    marginTop: 6
  }} />
      </label>
      <svg viewBox={`0 0 ${W} ${H}`} style={{
    width: "100%",
    maxWidth: 560,
    display: "block",
    margin: "4px 0"
  }} role="img" aria-label={`Reading the top ${read}% finds ${d.protocol[step]}% with the lawyer's reading, ${d.request[step]}% with the request's words`}>
        <line x1={L} y1={yv(0)} x2={W - 4} y2={yv(0)} stroke={grey} />
        <line x1={L} y1={yv(80)} x2={W - 4} y2={yv(80)} stroke={grey} strokeDasharray="3 3" />
        <path d={path(d.request)} fill="none" stroke="rgba(128,128,128,0.6)" strokeWidth={2} />
        <path d={path(d.protocol)} fill="none" stroke={teal} strokeWidth={2.5} />
        <circle cx={x(kx)} cy={yv(d.keywords[1])} r={4.5} fill={red} />
        <line x1={x(step)} y1={yv(0)} x2={x(step)} y2={yv(100)} stroke={gold} strokeWidth={1.5} />
      </svg>
      <div style={{
    fontSize: 12,
    opacity: 0.7,
    display: "flex",
    justifyContent: "space-between"
  }}><span>read 0%</span><span>dashed line: 80% found</span><span>50%</span></div>
      <div style={{
    fontSize: 14,
    marginTop: 10,
    lineHeight: 1.7
  }}>
        <div><span style={{
    color: teal
  }}>━</span> With the lawyer's reading: <b>{d.protocol[step]}%</b> of what the lawyer wanted is in front of a person</div>
        <div><span style={{
    opacity: 0.6
  }}>━</span> With the request's words: <b>{d.request[step]}%</b></div>
        <div><span style={{
    color: red
  }}>●</span> Keyword search reads {d.keywords[0]}% and finds {d.keywords[1]}%</div>
      </div>
    </div>;
};

export const FoundBars = ({bars}) => {
  return <div className="not-prose" style={{
    margin: "16px 0",
    fontSize: 14
  }}>
      {bars.map(([label, pct], k) => <div key={label} style={{
    margin: "10px 0"
  }}>
          <div style={{
    display: "flex",
    justifyContent: "space-between",
    gap: 8
  }}>
            <span>{label}</span><span><b>{Math.round(pct)}</b> in 100 found</span>
          </div>
          <div style={{
    height: 14,
    marginTop: 4,
    background: "rgba(128,128,128,0.15)",
    borderRadius: 4
  }}>
            <div style={{
    width: `${pct}%`,
    minWidth: 3,
    height: 14,
    borderRadius: 4,
    background: k === bars.length - 1 ? "#7D969B" : "rgba(128,128,128,0.55)"
  }} />
          </div>
        </div>)}
    </div>;
};

In a lawsuit, the other side can ask for every document that "relates to onshore or offshore oil and gas drilling or extraction activities". At Enron that means 455,449 emails. Nobody can read them all. So which ones do you read?

<Info>
  Lawsuits, audits, investigations and journalists with a leaked archive all face this: a pile too big to read and a request about what's in it. The request is only words. What counts is how the lawyer in charge reads them, and hunch is for turning that reading into a question you can run, measure and change.
</Info>

The experiment runs on 25,353 judged emails sampled from the 455,449-email collection, then weights the results to estimate collection-wide rates. Its added instructions condense the review protocol behind the answer key. That gives this comparison an advantage over a reader who has only the request's wording.

## Try it first

In 2010 a senior lawyer played this role for a public exercise, the TREC Legal Track, and wrote down what the drilling request covers. Would the lawyer count these?

<ScopePuzzle
  items={[
{ doc: "A weekly internal report that mentions, among other things, the company's drilling plans", yes: true, why: "Any document that discusses the company's own drilling or extraction counts, even in passing." },
{ doc: "A newspaper summary about the oil industry, emailed automatically every morning", yes: false, why: "News about the industry doesn't count, unless an employee forwards it with comments tying it to the company." },
{ doc: "A gas purchase agreement", yes: true, why: "Agreements that earn revenue from the business count: gas purchase, sale, gathering and field services." },
{ doc: "Emails about moving gas through the company's pipelines", yes: true, why: "The lawyer reads extraction as including transport by pipeline." },
{ doc: "A memo on staffing and headcount in the drilling business", yes: false, why: "Staffing documents are not about the activity itself." },
{ doc: "Sections of federal regulations on offshore drilling", yes: false, why: "Regulations, and testimony about the industry, are about drilling in general, not the company's." },
]}
/>

The words say drilling. The lawyer meant the whole business of getting oil and gas out of the ground, moved and paid for, as the company did it, and not the talk around it.

## 1. Get the emails

```bash theme={null}
D=prototype/examples/discovery
uv run python $D/fetch.py
```

This downloads the Enron emails and the judgments. Professional reviewers read 25,353 emails, not 455,449, and the senior lawyers settled the disputed calls: a sample, drawn so that each email stands for others, from one to about 150. It works like an opinion poll, where each person asked speaks for thousands. A spec's `weights:` tells hunch how many each one speaks for, so every number on this page is about the whole collection.

## 2. Ask the request, word for word

```yaml prototype/examples/discovery/request_only/drilling.yml theme={null}
judgment: drilling
source: ../.cache/drilling.csv
key: id
state: [email, attachments]
questions:
  drilling:
    type: noul      # yes or no, with how sure
    instructions: >-
      `email` is an email from Enron's
      files ... the other side asked Enron
      to produce "All documents or
      communications that describe,
      discuss, refer to, report on, or
      relate to onshore or offshore oil
      and gas drilling or extraction
      activities, ..." Does this email,
      with its attachments, fall under
      that request?
    gold: gold      # the lawyers' judgment
```

```bash theme={null}
hunch run $D/request_only/drilling.yml \
  --max-cost 0.5
```

It reads 5,837 emails in under two minutes for \$0.27, and finds 8 in every 100 of the emails the lawyer wanted. It read the request exactly as written, and that was the problem.

## 3. Write down what the lawyer meant

Add the lawyer's reading to the question, in a paragraph (condensed from eight pages of guidance):

```yaml prototype/examples/discovery/drilling.yml theme={null}
    instructions: >-
      ... The senior lawyer on the case
      reads the request broadly. It covers
      the company's own oil and gas
      drilling or extraction business ...
      and this includes moving oil or gas
      through pipelines. It also covers
      revenue from that business ...
      agreements that earn it ... Not
      responsive: news or third-party
      reports about the industry in
      general ...
```

<FoundBars bars={DISCOVERY_FOUND} />

Same model, same emails. The only change is a paragraph that says what the request means.

Before trusting the new wording, `hunch diff` asks it of every email and compares:

```bash theme={null}
hunch diff $D/drilling.yml \
  --against $D/request_only/drilling.yml
```

```text theme={null}
drilling: 869/5837 rows flip
  (134 within noise band)
  gold accuracy on the 5837 shared rows
  with gold 86.2% → 88.1%
  (✓ 489 fixed, ✗ 380 broken, ...)
  paired sign test p=0.000 → significant
```

489 fixed and 380 broken sounds like a small win. Look at which. Of the fixes, 439 are emails the lawyer wanted that the plain question had missed; of the breaks, 367 are emails nobody wanted, now flagged for a person to read and set aside. In a review a miss is the expensive mistake. Weighted to the whole collection, the share of wanted emails labelled "yes" rises from 8 in 100 to 63. That is the classifier's recall at its usual yes/no cutoff, before choosing how much a person will read.

## 4. Decide how much a person reads

Nobody hands over what a model picked without reading it. The model's job is the other side: setting aside the emails it is sure are not wanted, so a person reads the rest. `act` says how sure a "no" must be to be set aside, and `min_recall` says how much of what the lawyer wanted must still reach a person:

```yaml prototype/examples/discovery/privileged.yml theme={null}
    act: {yes: 1.0, no: 0.75}
    # a person reads every yes and every
    # unsure no; a sure no is set aside
tests:
  privileged: {min_recall: 0.8}
```

This is the privilege review, the emails Enron could keep from the other side because they carry legal advice. `hunch test` checks the worst plausible case, the bottom of the 95% interval, not the average:

```text theme={null}
PASS recall of yes 84.1%
  (95% CI 80.5%–87.2%)
  of 941 gold-yes rows at act 0.75:
  a person reads 15% of rows, the
  confident "no" answers are set aside
  (min 80% on the interval's lower bound)
```

A person reads about 68,000 emails instead of 455,449, and even at the bottom of the interval, 80% of what should be withheld is among them. This is a different measure from the 63% above: a person reads every "yes" and the uncertain "no" answers, so the reading queue can recover wanted emails the plain classifier labelled "no". That bar is the example's; a real privilege review, where every privileged email handed over can waive the protection, would set it much higher and read more. Drag the line to see what each amount of reading buys:

<ReadSlider data={DISCOVERY_READ} />

The red dot is a keyword search, a broad OR query of the kind a reviewer tries first (the exact queries are under "How it was measured"). On privilege and lobbying, reading the same number of emails in hunch's order finds more: 84 in 100 privileged emails against the keywords' 67. On drilling the keywords win, 91 against 84, because drilling is a topic of words: *pipeline*, *reserves*, *exploration*. Even there the winning query needs *pipeline*, which only the lawyer's reading puts in scope. Privilege is a question of who wrote to whom and why, and no list of words captures that.

## The lesson

The teams in the 2010 exercise, e-discovery firms and universities with up to ten hours of the same lawyer's time, hit the same wall on drilling. None found more than 25% of what the lawyer wanted, and the organisers wrote that none captured the lawyer's broad reading of the request.

A request is not a question until someone writes down what it means. In hunch that writing is the spec: it is read, diffed and tested like code, so when the lawyer's reading changes you see which documents move before anyone relies on them.

## Use it on your data

Point `source:` at your documents, put the request and your review protocol in `instructions`, have a lawyer judge a few hundred, and `hunch test` tells you how much a person must read.

## How it was measured

<AccordionGroup>
  <Accordion title="The data">
    The [TREC 2010 Legal Track](https://trec.nist.gov/data/legal10.html) Interactive task: four requests for production from a mock complaint (drilling, responses to spills, lobbying, and a privilege review), over the [EDRM Enron Email Data Set v2](https://trec-legal.umiacs.umd.edu/corpora/trec/legal10/) (ZL Technologies, CC BY 3.0 US), 455,449 messages. Professional review firms judged a stratified sample of each topic; the topic's senior lawyer (the "Topic Authority") then adjudicated about 10% of those calls, the ones teams appealed plus a sample of the rest. This page uses the final judgments, 25,353 in all, leaving out 154 marked unreadable. Each judgment carries the probability its email was sampled; summing one over it gives the collection's 455,449 for every topic (within 0.3% once the unreadable ones are left out), and `weights:` in each spec reweights by it. `fetch.py` keeps the first 20,000 characters of an email and 10,000 of its attachments; the specs send 6,000 and 2,000. The emails name real people; they are public records, and this page quotes no email and names no one.
  </Accordion>

  <Accordion title="The results, in full">
    From `measure.py`, weighted to the collection, 95% intervals from 1,000 bootstrap resamples within strata; hunch says yes at p(yes) ≥ 0.5.

    | Topic | Reader | Found (recall) | Right when it says yes (precision) | F1 |
    | - | - | - | - | - |
    | Drilling | keywords | 91% (87–94) | 22% (20–25) | 36% |
    | | request's words | 8% (6–10) | 34% (24–45) | 13% |
    | | lawyer's reading | 63% (57–69) | 46% (41–52) | 53% (48–58) |
    | Privileged | keywords | 67% (62–72) | 21% (19–23) | 32% |
    | | request's words | 25% (22–28) | 53% (47–59) | 34% |
    | | lawyer's reading | 44% (40–49) | 53% (48–58) | 48% (44–52) |
    | Lobbying | keywords | 77% (72–82) | 16% (14–17) | 26% |
    | | request's words | 77% (73–81) | 54% (49–59) | 64% |
    | | lawyer's reading | 68% (64–73) | 61% (55–66) | 64% (60–68) |

    Keywords, as regular expressions on the lowercased email and attachments: drilling `drill|\boil\b|\bgas\b|extraction|pipeline|reserves|exploration`; privilege `privilege|attorney|counsel|lawyer|legal|litigation|lawsuit|confidential`; lobbying `lobby|legislat|senator|congress|regulat|governor|\bbill\b|testimony|government affairs`. They read 17%, 14% and 13% of the collection; hunch's ranking, reading as much, finds 84%, 84% and 94%. (A first, narrower query, `drill|oil and gas|\bwell\b|\brig\b`, found 19% on drilling: a keyword search is only as good as its words.) On lobbying the request's words were already enough; the lawyer's reading moved precision up and recall down. To find 80% by ranking alone, a person reads the top 12.7% (9.3–17.4) for drilling, 12.2% (10.4–14.7) for privilege, 4.5% (3.8–5.1) for lobbying. The `min_recall` checks use a Wilson interval on the effective sample size (941 privileged emails weigh as 453).
  </Accordion>

  <Accordion title="The recall bars were chosen on the same answers">
    Each spec's `act` for "no" (0.82 drilling, 0.75 privilege, 0.65 lobbying) is the lowest that `hunch test` reported as keeping recall at 80% on the lower bound, found on the same judged emails it is then checked against, so the PASS is optimistic. A fair check: choose the bar on half the emails (split by a hash of the id) and test it on the other half. Chosen on one half: 0.86, 0.77, 0.65; on the other half recall is 89.9% (83.3–94.1), 87.9% (83.2–91.4) and 83.3% (78.5–87.2). Drilling and privilege hold; lobbying's lower bound falls just under 80%. On your own documents, choose the bar on one set of reviewed emails and check it on another.
  </Accordion>

  <Accordion title="Why spills is left out">
    Only 184 of the judged emails were about responses to spills, and three of them, sampled from the part of the collection no team flagged, stand for most of the 575 spill emails the sample implies are in the collection. Weighted, the 184 count as 8. hunch's recall there is somewhere between 9% and 44%, and `hunch test` says no recall bar can be promised; the spec is in the example without one.
  </Accordion>

  <Accordion title="The human reviewers">
    The first-pass reviewers' calls are published too, and against the final judgments they score far higher (F1 83% on drilling, 87% on privilege). That comparison is largely circular: the final judgments are the reviewers' calls, and only about 10% of them went to the senior lawyer (appeals, plus 884 others the organisers chose, of which about a quarter were overturned). Everywhere else the reviewer's call is the answer key. It is an upper bound on the reviewers, not a measurement. For a fair comparison of people and software on the 2009 edition of this exercise, see Grossman and Cormack, "Technology-Assisted Review in E-Discovery Can Be More Effective and More Efficient Than Exhaustive Manual Review", Richmond Journal of Law and Technology, 2011.
  </Accordion>

  <Accordion title="Against the 2010 teams, and why it isn't a race">
    The track's overview reports each team's final result: the best F1 was 26% on drilling (no team above 25% recall), 67% on lobbying and 41% on privilege, against hunch's 53%, 64% and 48% above. It is not a fair race. The paragraph each spec adds is our condensation of the instructions the reviewers who produced the judgments worked from, compiled largely from what each lawyer told the teams (the published version is dated after the teams had submitted). That is the answer key's own instrument, a real advantage the teams didn't have in that form; they questioned the lawyer directly, some for up to ten hours, some not at all. The lobbying paragraph also keeps two names specific to this case from those instructions (the Independent Energy Producers Association and the CPUC). A 2026 model is being compared with 2010 systems, and the Enron emails are all over the web (the judgments much less so). What the comparison does show is the cost: about \$0.33 and under two minutes per topic.
  </Accordion>

  <Accordion title="Cost">
    Both wordings of all four topics, every judged email: $2.54 in total, including a 400-email trial run. The answers are kept, so every number here, `measure.py` and the widgets are rebuilt for $0.
  </Accordion>
</AccordionGroup>
