> ## Documentation Index
> Fetch the complete documentation index at: https://fuguai.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Read an NDA before a lawyer does

> Ask real non-disclosure agreements the questions a lawyer checks first, and see why the question must ask what a contract says, not what it implies.

export const NDA_CLAUSES = [["must_be_marked", 25, 44, 61], ["advisors_allowed", 51, 52, 61], ["copies_allowed", 21, 30, 61], ["may_keep_after_return", 35, 41, 61], ["survives_termination", 44, 46, 61]];

export const NDA_FIX = [["says_otherwise", "not_mentioned"], ["not_mentioned", "not_mentioned"], ["not_mentioned", "not_mentioned"], ["says_otherwise", "says_otherwise"], ["says_otherwise", "says_otherwise"], ["says_otherwise", "says_otherwise"], ["says_otherwise", "not_mentioned"], ["says_otherwise", "not_mentioned"], ["says_otherwise", "not_mentioned"], ["says_otherwise", "not_mentioned"], ["says_otherwise", "says_otherwise"], ["says_otherwise", "not_mentioned"], ["says_otherwise", "not_mentioned"], ["not_mentioned", "not_mentioned"], ["says_otherwise", "not_mentioned"], ["says_otherwise", "says_otherwise"], ["says_otherwise", "not_mentioned"], ["says_otherwise", "says_otherwise"], ["says_otherwise", "not_mentioned"], ["says_otherwise", "not_mentioned"], ["says_otherwise", "not_mentioned"], ["not_mentioned", "not_mentioned"], ["says_otherwise", "not_mentioned"], ["says_otherwise", "not_mentioned"], ["not_mentioned", "not_mentioned"], ["says_otherwise", "not_mentioned"], ["says_otherwise", "not_mentioned"], ["says_otherwise", "not_mentioned"], ["says_otherwise", "says_otherwise"], ["says_otherwise", "says_otherwise"], ["says_otherwise", "not_mentioned"], ["says_otherwise", "not_mentioned"]];

export const NDA_PUZZLES = [{
  "id": "181",
  "chars": 8248,
  "excerpt": "“Confidential Information” shall mean all information in whatever form, whether imparted orally or in writing or by other medium including all copies of the same which one party hereto discloses to the other pursuant to the Purpose.",
  "elsewhere": 0,
  "why": "It covers everything, so a reader infers unmarked information counts. The contract never says so, and that is what the lawyers label.",
  "gold": "not_mentioned",
  "hunch": ["not_mentioned", {
    "says_so": 0.0,
    "says_otherwise": 0.29,
    "not_mentioned": 0.71
  }],
  "first": ["says_otherwise", {
    "says_so": 0.0,
    "says_otherwise": 0.76,
    "not_mentioned": 0.24
  }]
}, {
  "id": "117",
  "chars": 5099,
  "excerpt": "For convenience, the Disclosing Party may, but is not required to, mark written Confidential Information with the legend \"Confidential\" or an equivalent designation.",
  "elsewhere": 0,
  "why": "Marking is allowed but not required, so unmarked information is protected, and the contract puts that in words.",
  "gold": "says_otherwise",
  "hunch": ["says_otherwise", {
    "not_mentioned": 0.01,
    "says_so": 0.0,
    "says_otherwise": 0.99
  }],
  "first": ["says_otherwise", {
    "says_otherwise": 1.0,
    "not_mentioned": 0.0,
    "says_so": 0.0
  }]
}, {
  "id": "157",
  "chars": 5810,
  "excerpt": "information disclosed by the Disclosing Party will be considered Confidential Information by the Receiving Party only if such information is conspicuously designated as \"Confidential\" (i) in writing, if communicated in writing, or (ii) confirmed in writing within thirty (30) days of disclosure, if disclosed orally or in other non-tangible form",
  "elsewhere": 0,
  "why": "Only information designated \"Confidential\", or confirmed in writing, counts.",
  "gold": "says_so",
  "hunch": ["says_so", {
    "not_mentioned": 0.0,
    "says_otherwise": 0.0,
    "says_so": 1.0
  }],
  "first": ["says_so", {
    "says_so": 1.0,
    "says_otherwise": 0.0,
    "not_mentioned": 0.0
  }]
}, {
  "id": "51",
  "chars": 6016,
  "excerpt": "All information disclosed by a Party or by Affiliates of a Party to the other Party or its respective Affiliates orally, electronically, writing or by any other means during the data sharing negotiations shall be considered as confidential unless expressly stated otherwise by the disclosing Party.",
  "elsewhere": 0,
  "why": "Everything counts unless the discloser says otherwise. hunch reads that as an answer, and is sure of it; the lawyers' label says the contract never mentions marking.",
  "gold": "not_mentioned",
  "hunch": ["says_otherwise", {
    "says_so": 0.0,
    "not_mentioned": 0.0,
    "says_otherwise": 1.0
  }],
  "first": ["says_otherwise", {
    "says_so": 0.01,
    "says_otherwise": 0.99,
    "not_mentioned": 0.0
  }]
}];

export const ClauseBars = ({clauses}) => {
  const NAMES = {
    must_be_marked: "Must it be marked?",
    advisors_allowed: "May advisors see it?",
    copies_allowed: "May it be copied?",
    may_keep_after_return: "May a copy be kept after return?",
    survives_termination: "Do duties outlive the agreement?"
  };
  const teal = "#7D969B";
  const bar = (n, of, bg) => <div style={{
    height: 10,
    marginTop: 3,
    background: "rgba(128,128,128,0.15)",
    borderRadius: 4
  }}>
      <div style={{
    width: `${100 * n / of}%`,
    height: 10,
    background: bg,
    borderRadius: 4
  }} />
    </div>;
  return <div className="not-prose" style={{
    margin: "16px 0",
    fontSize: 14
  }}>
      {clauses.map(([q, first, now, of]) => <div key={q} style={{
    margin: "12px 0"
  }}>
          <div style={{
    display: "flex",
    justifyContent: "space-between",
    gap: 8,
    flexWrap: "wrap"
  }}>
            <span>{NAMES[q]}</span>
            <span style={{
    fontVariantNumeric: "tabular-nums"
  }}><b>{now}</b> of {of} right <span style={{
    opacity: 0.7
  }}>(was {first})</span></span>
          </div>
          {bar(first, of, "rgba(128,128,128,0.45)")}
          {bar(now, of, teal)}
        </div>)}
      <div style={{
    fontSize: 13,
    opacity: 0.75
  }}>
        <span style={{
    color: "rgba(128,128,128,0.8)"
  }}>■</span> first wording · <span style={{
    color: teal
  }}>■</span> reworded
      </div>
    </div>;
};

export const FixDots = ({fix}) => {
  const teal = "#7D969B", red = "#C64D35", gold = "#C9A227";
  const color = {
    not_mentioned: teal,
    says_otherwise: red,
    says_so: gold
  };
  const rows = [["First wording", 0], ["Reworded", 1]];
  return <div className="not-prose" style={{
    margin: "16px 0",
    fontSize: 14
  }}>
      {rows.map(([title, k]) => {
    const wrong = fix.filter(f => f[k] === "says_otherwise").length;
    return <div key={title} style={{
      margin: "10px 0"
    }}>
            <div style={{
      display: "flex",
      justifyContent: "space-between",
      gap: 8,
      flexWrap: "wrap"
    }}>
              <span>{title}</span><span><b>{wrong}</b> of {fix.length} read as “unmarked counts too”</span>
            </div>
            <div style={{
      display: "flex",
      flexWrap: "wrap",
      gap: 4,
      marginTop: 6
    }} role="img" aria-label={`${title}: ${wrong} of ${fix.length} contracts answered "unmarked counts too"`}>
              {fix.map((f, i) => <span key={i} style={{
      width: 12,
      height: 12,
      borderRadius: 6,
      background: color[f[k]]
    }} />)}
            </div>
          </div>;
  })}
      <div style={{
    fontSize: 13,
    opacity: 0.75,
    marginTop: 6
  }}>
        <span style={{
    color: teal
  }}>●</span> doesn't say (the lawyers' label) · <span style={{
    color: red
  }}>●</span> unmarked counts too
        · <span style={{
    color: gold
  }}>●</span> must be marked
      </div>
    </div>;
};

export const ClausePuzzle = ({puzzles}) => {
  const [i, setI] = useState(0);
  const [pick, setPick] = useState(null);
  const z = puzzles[i];
  const shown = pick !== null;
  const teal = "#7D969B", red = "#C64D35", grey = "rgba(128,128,128,0.25)";
  const OPTS = [["says_so", "Yes, it must be marked"], ["says_otherwise", "No, unmarked counts too"], ["not_mentioned", "It doesn't say"]];
  const name = l => OPTS.find(o => o[0] === l)[1].toLowerCase();
  const [label, probs] = z.hunch;
  const btn = {
    border: `1px solid ${grey}`,
    borderRadius: 8,
    padding: "7px 10px",
    background: "transparent",
    color: "inherit",
    font: "inherit",
    fontSize: 14,
    textAlign: "left"
  };
  return <div className="not-prose hunch-widget" style={{
    border: `1px solid ${grey}`,
    borderRadius: 12,
    padding: 16,
    margin: "16px 0"
  }}>
      <style>{`.hunch-bar{transition:width .4s ease-out} @media (prefers-reduced-motion: reduce){.hunch-bar{transition:none}}`}</style>
      <div style={{
    fontSize: 13,
    opacity: 0.7
  }}>
        Contract {i + 1} of {puzzles.length}, {z.chars.toLocaleString("en")} characters long. The sentence that matters:
      </div>
      <blockquote style={{
    margin: "8px 0",
    padding: "8px 12px",
    borderLeft: `3px solid ${teal}`,
    fontSize: 15,
    lineHeight: 1.6,
    overflowWrap: "anywhere"
  }}>
        {z.excerpt}
      </blockquote>
      {z.elsewhere === 0 && <div style={{
    fontSize: 13,
    opacity: 0.7
  }}>Nowhere else does it mention marking, labels or designations.</div>}
      <div style={{
    fontSize: 15,
    fontWeight: 600,
    margin: "10px 0 6px"
  }}>Must information be marked confidential to count?</div>
      <div style={{
    display: "flex",
    flexWrap: "wrap",
    gap: 8
  }}>
        {OPTS.map(([l, text]) => <button key={l} onClick={() => !shown && setPick(l)} disabled={shown} style={{
    ...btn,
    flex: "1 1 160px",
    cursor: shown ? "default" : "pointer",
    borderColor: shown && l === z.gold ? teal : pick === l ? "currentColor" : grey
  }}>
            {text}{shown && l === z.gold ? " ✓" : ""}
            {shown && pick === l && <span style={{
    opacity: 0.7,
    fontSize: 13
  }}> · your answer</span>}
          </button>)}
      </div>
      {shown && <div style={{
    marginTop: 12
  }}>
          <div style={{
    fontSize: 13,
    opacity: 0.8,
    marginBottom: 2
  }}>hunch, and how sure it was of each answer:</div>
          {OPTS.map(([l, text]) => <div key={l} style={{
    display: "flex",
    alignItems: "center",
    gap: 8,
    margin: "4px 0",
    fontSize: 13
  }}>
              <span style={{
    width: 132,
    flexShrink: 0,
    fontWeight: l === label ? 600 : 400
  }}>{text}</span>
              <div style={{
    flex: 1,
    height: 6,
    background: grey,
    borderRadius: 3
  }}>
                <div className="hunch-bar" style={{
    width: `${Math.round(probs[l] * 100)}%`,
    height: 6,
    borderRadius: 3,
    background: l !== label ? "rgba(128,128,128,0.55)" : label === z.gold ? teal : red
  }} />
              </div>
              <span style={{
    minWidth: 34,
    textAlign: "right",
    fontVariantNumeric: "tabular-nums"
  }}>{probs[l].toFixed(2)}</span>
            </div>)}
          <p style={{
    fontSize: 14,
    margin: "10px 0 0",
    lineHeight: 1.6
  }}>
            <b style={{
    color: pick === z.gold ? teal : red
  }}>{pick === z.gold ? "You agree with the lawyers." : "Your answer isn't the lawyers'."}</b>{" "}
            They labelled it “{name(z.gold)}”. hunch said “{name(label)}”{label === z.gold ? ", the same as them." : ", which is not."}{" "}
            {z.first[0] !== label ? `The first wording of the question said “${name(z.first[0])}”. ` : ""}{z.why}
          </p>
        </div>}
      <div style={{
    display: "flex",
    justifyContent: "flex-end",
    marginTop: 10
  }}>
        <button onClick={() => {
    setI((i + 1) % puzzles.length);
    setPick(null);
  }} style={{
    ...btn,
    padding: "6px 12px",
    cursor: "pointer"
  }}>
          Next contract →
        </button>
      </div>
    </div>;
};

This is how one real non-disclosure agreement defines its secrets:

> “Confidential Information” shall mean all information in whatever form, whether imparted orally or in writing or by other medium including all copies of the same which one party hereto discloses to the other pursuant to the Purpose.

A lawyer reviewing it asks: must information be marked confidential to count? What does this contract say?

<Info>
  A company selling a business sends an NDA to every bidder, and each comes back on the bidder's own template. The same few questions get asked of every one. Finding the answer in a contract is a judgment, and that is the part hunch is for. What to do about the answer stays with a lawyer.
</Info>

## Try it first

Read the sentence and pick an answer. Then see what the lawyers who labelled the contract said, and what hunch said and how sure it was.

<ClausePuzzle puzzles={NDA_PUZZLES} />

The first contract is the one above. It covers all information, so it is tempting to answer that unmarked information counts. But the contract never says so; the reader infers it. The lawyers who labelled these contracts count only what a contract says in words, and a question checked against their labels has to do the same.

## 1. Collect contracts with known answers

[ContractNLI](https://stanfordnlp.github.io/contract-nli/) is a set of 607 NDAs collected from the web. Lawyers labelled each against 17 fixed statements, such as "information must be marked to count", with one of three answers: the contract says so, says the opposite, or doesn't mention it.

The example takes five of those statements and the 123 contracts of the set's test split, one row per contract:

```text contractnli_test.csv theme={null}
id, file_name, text,
gold_must_be_marked, gold_advisors_allowed,
gold_copies_allowed, gold_may_keep_after_return,
gold_survives_termination
```

`text` is the whole contract; the median one is about 9,600 characters. The `gold_` columns are the lawyers' labels. hunch calls a known answer like that gold, and checks its own answers against it.

## 2. Write the questions

A spec is one file that says what to read and what to ask. Here the model reads the whole contract and answers five questions. Each is a choice between three answers, and the criteria say what each answer means:

```yaml nda_review.yml theme={null}
# one question shown, long lines folded: same text
judgment: nda_review
model: jev-1.13.0
source: contractnli_test.csv
key: id
state: [text]           # what the model reads
clip: {text: 60000}
questions:
  must_be_marked:
    type: choice
    instructions: >-
      In this non-disclosure agreement (`text`),
      must information be expressly identified or
      marked as confidential by the disclosing party
      to count as Confidential Information?
    criteria:
      says_so: >-
        The agreement expressly says only information
        that is marked, labelled or expressly identified
        as confidential (or confirmed as such in
        writing) is protected
      says_otherwise: >-
        The agreement expressly says information is
        protected even when it is not marked or
        identified (e.g. "whether or not marked", oral
        information not confirmed in writing)
      not_mentioned: >-
        The agreement doesn't expressly say either way.
        A broad definition of Confidential Information
        alone is not_mentioned
    gold: gold_must_be_marked   # the lawyers' label
  # four more questions of the same shape:
  # advisors_allowed, copies_allowed,
  # may_keep_after_return, survives_termination
```

## 3. Run it

```sh theme={null}
hunch run prototype/examples/nda_review/nda_review.yml \
  --max-cost 0.05
```

The five questions share one request per contract. Asking them of all 123 contracts cost \$0.016 the first time; every answer is kept, so running it again is free. Each answer is one of the three labels and a probability: how sure the model is of it.

## 4. Check it against the lawyers

```sh theme={null}
hunch test prototype/examples/nda_review/nda_review.yml
```

```text theme={null}
must_be_marked (choice, 123 rows)
  gold: 123 rows (123 from source, 0 from review)
  PASS accuracy 82.9% (min 0%)
```

That number is on the contracts the question was written from, which flatters it. The fair measure is 61 contracts it has never seen.

## The result

On 61 fresh contracts, how many answers to each question agree with the lawyers:

<ClauseBars clauses={NDA_CLAUSES} />

Rewording helped most on marking, and copying is still right only about half the time.

## One word

The first wording of the questions asked what a contract means. Its criteria read "information is protected even when it is not marked", with no word about where that has to be written. On the marking question it agreed with the lawyers on barely half the contracts, and it was sure of its mistakes.

The mistakes had one shape. The lawyers labelled 63 of the contracts silent on marking, and the model said "unmarked counts too" for 43 of them. It had done what you may have done on the first contract: read a broad definition and inferred the rest. The same shape was the most common mistake in all five questions. A duty to return "all information" was read as "no copies may be kept", and a five-year confidentiality period as "the obligations outlive the agreement".

The model was answering a slightly different question from the one the labels answer: what does this contract imply, rather than what does it say. The fix was in the criteria. They now ask what the agreement *expressly* says, and each "not mentioned" names the inference to avoid, such as "a broad definition of Confidential Information alone is not\_mentioned".

Before adopting it, `hunch diff` asked the new wording of every contract and compared:

```text theme={null}
# trimmed and wrapped
must_be_marked: 45/123 rows flip
  gold accuracy on the 123 shared rows with gold
    52.8% → 82.9%  (✓ 41 fixed, ✗ 4 broken)
  paired sign test p=0.000 → significant
```

But those are the contracts whose mistakes the new wording was written from. So here are the 32 fresh contracts the lawyers called silent on marking, one dot each, as each wording answered them:

<FixDots fix={NDA_FIX} />

A question can be precise and still ask the wrong thing. When the mistakes share one shape, read them before adding rules, and test the fix on rows it was not written from.

## Where it stands

This is a first pass, not a decision. The marking, advisors and survival questions agree with the lawyers about three times in four or better. The copying question agrees half the time, and a third of the copying answers it gave at 0.9 or above were wrong. Software should not act on that column.

It can sort a pile. The contracts that clearly let advisors see the data, or clearly keep obligations alive, go to the bottom of a lawyer's stack, and the rest get read first.

## Use it on your data

Put your contracts in a CSV with an `id` and a `text` column, point `source:` at it and run the spec. Without lawyers' labels there is nothing to test against, so [review](/guides/review) a few dozen contracts first; `hunch test` then tells you how often it agrees with your reviewers.

## How it was measured

<AccordionGroup>
  <Accordion title="The data">
    [ContractNLI](https://stanfordnlp.github.io/contract-nli/) (Hitachi America, Ltd., CC BY 4.0): 607 NDAs collected from the web, each labelled by lawyers against 17 fixed hypotheses. `contractnli_test.csv` is its test split (123 NDAs), `contractnli_dev.csv` its dev split (61 NDAs). The five `gold_*` columns are hypotheses nda-1 (must\_be\_marked), nda-7 (advisors\_allowed), nda-17 (copies\_allowed), nda-20 (may\_keep\_after\_return) and nda-19 (survives\_termination), mapped Entailment → says\_so, Contradiction → says\_otherwise, NotMentioned → not\_mentioned (`NOTICE.md`). The median test contract is 9,614 characters and the longest in either split 41,779, so `clip: {text: 60000}` cuts none. The whole text is the state, and the five questions share one request per contract. Asking the first wording of all 123 test contracts cost \$0.016 (610 answers; 5 were already stored).
  </Accordion>

  <Accordion title="The results, in full">
    On the 123 test contracts, first wording → reworded: must\_be\_marked 52.8% → 82.9%, advisors\_allowed 68.3% → 74.8%, copies\_allowed 42.3% → 56.1%, may\_keep\_after\_return 67.5% → 68.3%, survives\_termination 74.0% → 78.9%. By `hunch diff` on test: marking 41 fixed, 4 broken (p=0.000); advisors 8 fixed, 0 broken (p=0.008); copies 18 fixed, 1 broken (p=0.000); keeping a copy 5 fixed, 4 broken (p=1.000); survival 8 fixed, 2 broken (p=0.109).

    On the 61 dev contracts, which played no part in the rewording:

    | Clause                | First wording | Reworded | Rows fixed / broken | Sign test |
    | --------------------- | ------------- | -------- | ------------------- | --------- |
    | must be marked        | 41.0%         | 72.1%    | 21 / 2              | p=0.000   |
    | may go to advisors    | 83.6%         | 85.2%    | 1 / 0               | p=1.000   |
    | copies allowed        | 34.4%         | 49.2%    | 9 / 0               | p=0.004   |
    | may keep after return | 57.4%         | 67.2%    | 8 / 2               | p=0.109   |
    | survives termination  | 72.1%         | 75.4%    | 3 / 1               | p=0.625   |

    The gain on marking holds (p \< 0.001 on dev), and so does copying's; the other three changes are within noise on 61 contracts, and nothing got worse in a way the test can see. On dev's 32 contracts labelled silent on marking, the first wording answered "says otherwise" for 27 and the reworded one for 8.
  </Accordion>

  <Accordion title="How sure it is, and how often that is right">
    From `hunch test` on the 123 test contracts, the share wrong among answers given at 0.9 or above, first wording → reworded: must\_be\_marked 27.1% → 11.3%, advisors\_allowed 18.0% → 15.4%, copies\_allowed 46.9% → 35.4%, may\_keep\_after\_return 14.5% → 9.0%, survives\_termination 15.9% → 4.9%. With the first wording, every question's most common mistake was a contract the lawyers labelled not\_mentioned, answered as if it spoke: 43 of those on marking, 14 on advisors, 47 on copies, 37 on keeping a copy, 25 on survival.
  </Accordion>

  <Accordion title="The question was reworded on the contracts it is scored on">
    The new wording was written after reading the test contracts' mistakes, so its test numbers are flattering; the dev numbers are the honest measure. The first wording is kept as `first/nda_review.yml`, so both comparisons can be run again for free from `prototype/examples/nda_review/`: `hunch diff nda_review.yml --against first/nda_review.yml` on test, and the same with `--source contractnli_dev.csv` on dev.
  </Accordion>

  <Accordion title="What does not work yet">
    Copying fails because permission to copy usually sits inside an exception to another clause ("except as necessary for the Purpose"), which is the "negation by exception" the dataset's authors single out as hard. After rewording, its most common mistakes on test are still answers of "copies allowed": 32 on contracts the lawyers labelled not\_mentioned and 17 on ones they labelled says\_otherwise. What would move it further is not more wording. Two ways to try, each a spec change `diff` can judge: ask a narrower question per clause (is copying mentioned at all, then is it allowed), or put the sentences that mention copies in the state instead of the whole contract.

    Some confident mistakes are close readings. The fourth contract in "Try it first" treats everything as confidential "unless expressly stated otherwise by the disclosing Party"; the lawyers labelled it silent on marking, and hunch answers "unmarked counts too" at 1.00.
  </Accordion>
</AccordionGroup>
