A general counsel I supported used four different outside counsel firms depending on deal type and jurisdiction. When she compared the indemnification language across five recent vendor agreements, all drafted from what was supposed to be the same internal playbook, she found five meaningfully different versions of the same clause, each firm having applied its own house style and risk tolerance despite being handed the same starting template and the same negotiating parameters. Two of the five had quietly narrowed the company's indemnification protection in ways that hadn't been flagged as a deviation to anyone internally.
Most legal departments that use multiple outside firms have some version of a playbook: standard positions on liability caps, indemnification scope, IP ownership, termination rights, the terms the company wants to hold firm on and the terms where some negotiation flexibility is acceptable. The playbook exists as a document, usually a PDF or Word file that gets emailed to outside counsel at the start of an engagement, and from that point forward, enforcement of the playbook depends entirely on individual attorneys remembering it and applying it consistently, which across four firms and dozens of attorneys is not a realistic expectation.
The drift isn't usually intentional or negligent. It's what happens naturally when different attorneys, each with their own drafting habits built over years of practice at different firms, work from the same starting point but apply their own judgment on phrasing, structure, and risk framing. The problem is that this drift compounds over time and across deals in ways nobody's tracking, because the in-house team is reviewing each contract individually, not comparing it systematically against the last twenty contracts of the same type.
The approach here isn't fundamentally different from clause extraction in due diligence, but the goal is different: instead of flagging risk against a static risk taxonomy, the system flags deviation from the company's own historical drafting pattern and stated playbook positions. This means the baseline isn't a generic industry standard, it's the company's own approved template plus a corpus of its own previously negotiated and approved agreements of the same type.
I build this as a clause-level comparison, extracting each key clause category (indemnification, limitation of liability, IP assignment, termination, governing law) from an incoming draft and comparing it both against the playbook's stated position and against a semantic similarity baseline built from the company's library of previously approved agreements of that same contract type. A clause that deviates meaningfully from both the stated playbook and the historical pattern gets flagged for review, with the specific prior agreement or playbook section it deviates from attached to the flag, so the in-house reviewer isn't starting from scratch, they're looking at a direct comparison.
Not every deviation is a problem. Sometimes outside counsel narrows or expands a clause because the specific deal circumstances genuinely warrant it, a higher-risk counterparty, an unusual deal structure, a jurisdiction with specific requirements that don't fit the standard template. The consistency check isn't meant to force rigid uniformity across every contract regardless of context, it's meant to surface deviations so someone makes a conscious decision about them rather than the deviation happening silently and getting discovered, if ever, only when the clause actually matters during a dispute.
I build a required field into the flagging workflow: when an attorney or in-house reviewer confirms a flagged deviation was intentional, they log the reason, which then feeds back into the system as an accepted variant for that deal type or counterparty profile, rather than a repeated deviation that keeps getting flagged as if it were new every time. Over time, this builds a more nuanced set of acceptable variation patterns instead of a single rigid template, without losing the original discipline of catching drift before it becomes the default.
A mistake I've seen in early attempts at this is building the comparison logic around one firm's typical drafting patterns, usually because that firm handles the majority of the company's contract volume and their language became the de facto baseline by accident rather than by design. This breaks down the moment a different firm handles a deal, because their drafting style gets flagged as deviation across the board even when the substance is equivalent to what the primary firm would have produced, just phrased differently. The baseline needs to be built from the company's own approved playbook and prior negotiated outcomes, explicitly independent of any single firm's house style, so the consistency check works the same way regardless of which firm is drafting.
The value isn't catching every stylistic variation, most of those genuinely don't matter. It's catching the substantive drift, like the narrowed indemnification language in that first example, before it gets executed and becomes the company's actual legal exposure rather than a draft that could have been caught and corrected. For that general counsel, running this consistency check across the four firms she used didn't reduce her reliance on outside counsel's legal judgment. It gave her a systematic way to verify that judgment was landing where she'd actually asked it to, deal after deal, instead of finding out five contracts later that it hadn't been.