Are Legal AI Agents Ready for Real Legal Work?
Legal AI is moving beyond chatbots and into agents designed to carry out multi-step legal work. New benchmarks show that these systems can already perform meaningful parts of legal workflows — but they also reveal a large gap between doing pieces of legal work well and reliably completing an entire assignment without lawyer intervention.
The next major phase of legal AI is already underway. For the past few years, most lawyers experienced generative AI through a familiar interface: a chat box. You asked a question, uploaded a document or requested a draft, and the AI responded. The lawyer remained the orchestrator, deciding what to ask next, supplying additional information and moving the work from one step to another. That model is beginning to change.
Legal technology companies are increasingly building AI agents designed not merely to answer questions but to perform work. Give the system an objective, a collection of documents and access to the right tools, and the agent can potentially decide what steps are required, retrieve information, analyze documents, produce intermediate outputs, revise its work and eventually deliver something resembling a finished legal work product.
Harvey now allows firms to build workflow agents capable of working with documents and generating or updating Word files. Thomson Reuters has rebuilt CoCounsel Legal around an agentic architecture designed to plan, retrieve information, reason across multiple steps and work toward legal deliverables. Legora and other major legal AI companies are similarly moving from assistants toward agentic workflows. The legal AI market is increasingly selling execution rather than conversation.
But that creates a much more important question than whether an agent can draft a clause or summarize a contract: Are legal AI agents actually ready to perform real legal work? The emerging research suggests that the answer is both more encouraging and more complicated than the product demonstrations sometimes imply.
Agents are already useful. In some structured workflows, they can perform surprisingly large portions of work that previously required significant lawyer time. But the evidence also shows how quickly performance can deteriorate when a legal assignment becomes longer, interconnected and dependent on multiple pieces of information.
The dividing line is increasingly not AI versus lawyer. It is bounded delegation versus unsupervised autonomy.
What exactly is a legal AI agent?
The word agent is becoming one of the most overused terms in technology, so it helps to be precise. A conventional legal AI assistant is primarily reactive. A lawyer asks it to summarize an agreement, draft a clause, explain a provision or research a question. The system produces an answer and waits for another instruction.
An agent receives something closer to an objective. Instead of asking, “Summarize the change-of-control clause in this agreement,” a lawyer might ask an agent to review the contracts in a data room, identify agreements containing change-of-control or assignment restrictions, determine which may require consent in connection with the transaction and prepare a diligence report with citations to the relevant provisions.
Completing that assignment requires more than generating text. The system has to break the objective into steps, identify the relevant documents, determine what information matters, use available tools, keep track of intermediate findings, compare results across documents, identify missing information and assemble those findings into the requested work product.
The practical capabilities can include reviewing contract populations, extracting structured information, building diligence tables, conducting multi-stage legal research, comparing documents, drafting agreements, updating trackers, preparing chronologies, applying playbooks and producing first drafts of legal work product. That is why agentic AI has attracted so much attention. The economic opportunity is much larger than making lawyers slightly faster at prompting a chatbot. The ambition is to automate portions of the workflow itself.
The real problem with measuring legal AI
Until recently, much of the evidence around legal AI came from relatively narrow benchmarks. Can a model answer a legal multiple-choice question? Can it classify a clause? Can it identify whether a contract contains a particular provision? Can it answer a legal research question? Those tests are useful, but legal practice rarely arrives in neatly isolated questions.
A partner does not normally walk into an associate’s office and ask whether paragraph 8.2 contains an assignment restriction. The assignment is much more likely to be something like: review these agreements, work out which contracts might create a closing issue, tell me which consents we need, check whether there are exceptions and give me something I can discuss with the client tomorrow. That assignment contains multiple dependent steps.
And that matters because small errors compound. An agent might correctly understand the law but retrieve the wrong document. It might retrieve the right document but overlook an amendment. It might identify the relevant clause but misunderstand an exception. It might correctly identify nine material risks and miss the tenth. Each individual component can look impressive while the final work product remains unacceptable.
That is precisely the problem a new generation of legal AI benchmarks is trying to capture.
Harvey’s Legal Agent Benchmark changes the question
In May 2026, Harvey released the Legal Agent Benchmark, or LAB , an open-source benchmark specifically designed to evaluate agents on long-horizon legal work. Rather than asking models isolated legal questions, LAB attempts to mimic how assignments are actually delivered inside a law firm. An agent receives an instruction, a synthetic client matter containing relevant materials and a requirement to produce a work product.
The result is then assessed against an expert-written rubric covering the facts, analysis, format and other elements expected from the assignment. The benchmark includes more than 1,200 tasks across 24 practice areas and more than 75,000 expert-written rubric criteria. The important point is that the benchmark is trying to evaluate an agent in something resembling a legal working environment, rather than simply asking whether an LLM can answer a question correctly.
LAB also uses what Harvey calls an all-pass standard. For a task to count as successfully completed, every required rubric criterion must pass. That may sound harsh, but it reflects an uncomfortable reality about legal work.
If a diligence report correctly identifies eight material issues but misses the ninth issue that allows a major customer to terminate after closing, the report is not “89% successful” in any commercially meaningful sense. If a research memo correctly identifies four authorities but relies on a fifth case for a proposition the case does not support, the client cannot safely use 80% of the memo.
Legal work is often conjunctive: several things have to be right at the same time for the work product to be dependable.
The initial results were sobering
Harvey published its first LAB baseline results in May 2026. Under the strict all-pass methodology, the frontier models initially tested completed less than 10% of benchmark tasks end-to-end in aggregate. Importantly, the models often performed substantially better when individual rubric criteria were considered separately.
In other words, the systems could do large portions of an assignment correctly while still failing to produce a completely acceptable final work product. That distinction may be the most important thing lawyers need to understand about agentic AI right now.
An agent can be extremely useful without being reliable enough to operate independently. Those are not contradictory statements.
Suppose an agent correctly performs 45 of the 50 steps that would otherwise consume several hours of associate time. That could represent extraordinary productivity. But if the five missed steps include a material fact, an exception to a contractual restriction or an adverse authority, a lawyer still needs to identify the omissions before the work leaves the firm.
This is why the question “Can AI do legal work?” is becoming too crude. A better question is: How much of this legal assignment can safely be delegated to the agent before professional review is required?
The numbers are improving quickly
There is another reason not to draw an overly pessimistic conclusion from the initial results. The technology is moving very quickly. Publicly tracked results on the Harvey benchmark have improved since the initial May testing, and newer frontier models and specialized systems are completing substantially more long-horizon tasks under strict evaluation.
That does not mean we should translate any one percentage into a statement such as “AI can now do 25% of legal work.” Benchmarks measure particular task populations, under particular agent harnesses, with particular models and evaluation rules. What they do show is that end-to-end reliability is improving while remaining significantly harder than isolated legal reasoning.
Other benchmarks point in the same direction
Harvey is not alone in trying to measure long-horizon legal agents. HAQQ released its own Legal Agent Study in 2026, covering more than 1,300 long-horizon legal tasks across 24 practice areas and tens of thousands of rubric criteria. Like Harvey, it uses a strict all-pass approach.
HAQQ reported considerably higher completion rates for its own specialized legal agent than for the general frontier models it tested. Those percentages should not be directly compared with Harvey’s results because the datasets, agent harnesses, model versions and evaluation methodologies differ. HAQQ is also a legal AI vendor evaluating its own system, which should always be kept in mind when interpreting vendor-produced benchmarks.
But the broader lesson is consistent: specialization and agent architecture matter. Legal performance is not determined solely by the underlying foundation model. How the model is connected to tools, documents, retrieval, memory, workflows and evaluation can materially change the final result.
Why long-horizon legal work is so difficult
Legal work is unusually hostile to error accumulation. Consider a relatively ordinary M&A diligence assignment. An agent may need to determine which documents fall within scope, identify the operative agreement rather than a superseded version, connect amendments to the original contract, locate assignment and change-of-control provisions, interpret defined terms, determine whether exceptions apply, identify whether consent or merely notice is required, extract the relevant language, classify the issue consistently, cite the correct source and then decide whether the issue meets the agreed escalation threshold.
An agent can be highly capable at each individual step. But completing the workflow successfully requires all of those capabilities to remain aligned across many documents and many stages of reasoning. This is the long-horizon problem.
And legal matters can be much longer than a diligence exercise. A litigation strategy may evolve over months or years. A transaction can move from term sheet to diligence, negotiation, signing, regulatory approvals and closing. An employment dispute can begin with an internal investigation, move into correspondence and settlement discussions, and later become litigation.
The most valuable context in legal work often does not exist inside one document. It exists across the matter.
This is why memory and context matter
The next major challenge for legal agents may therefore be less about generating better paragraphs and more about maintaining reliable matter context. A lawyer working on a transaction remembers that the client refused to accept a particular risk three weeks ago. The lawyer knows that the chief financial officer cares about one provision more than another and that an apparently ordinary clause creates unusual risk because of something discovered elsewhere in diligence.
The lawyer may also remember what opposing counsel said on a call that was never captured in the contract, or know that the partner considers one issue commercially immaterial even though the drafting looks aggressive on paper. Much of this is tacit context.
An agent operating only on the documents it has been given may not have access to it. That is one reason legal AI companies are increasingly building matter workspaces, persistent memory, knowledge integrations and workflow systems rather than simply improving their chat interfaces.
The battle over agentic legal AI may ultimately be a battle over context architecture.
Where agents already make sense
The weakness of fully autonomous agents does not mean legal teams should wait. There are already categories of work where agentic systems make considerable sense, particularly where the work is bounded, repetitive, objectively verifiable and supported by clear source material.
Due diligence and large-scale contract review
This is perhaps the clearest current use case. The task can be tightly scoped, the documents are known, the questions can be defined and the desired output can be structured. Crucially, the agent’s findings can be checked against the underlying contracts. A legal team can ask an agent to review a document population for assignment provisions, change-of-control clauses, consent rights, termination provisions or other defined issues and populate a review table.
The lawyer does not need to trust the agent blindly. The lawyer needs the agent to reduce the amount of mechanical work required to reach the verification stage.
Information extraction
Extraction is particularly suitable where the expected answer has a defined structure. Dates, party names, contract values, renewal periods, notice requirements, governing law, termination rights and other defined contractual fields can increasingly be extracted at scale and then reviewed by a lawyer or legal operations professional.
Document classification
Determining whether documents belong to particular categories can also fit the agentic model well. The system can classify contracts, correspondence, pleadings, discovery materials or regulatory documents before a human reviewer deals with the smaller set requiring attention.
Playbook-based contract review
Agents are also becoming increasingly useful when the lawyer has already converted legal judgment into a playbook. If a company has established preferred positions, fallback positions and escalation rules for NDAs or vendor agreements, the agent has a much more constrained environment in which to operate.
The important point is that the playbook provides the judgment framework. The agent applies it.
Regulatory and compliance monitoring
Agents can monitor defined sources, identify developments, classify relevance and prepare initial summaries for lawyers. Again, this works best where the output is a trigger for lawyer attention rather than the final legal conclusion.
The common characteristic is not “simple work”
It would be a mistake to describe all of these tasks as low-value or simple. Due diligence can be complex. Contract review can involve important commercial consequences. Regulatory monitoring can involve sophisticated law. What makes these tasks suitable for agents is something different.
They can often be bounded. The organization can specify what information the agent may use, what question it must answer, what tools it may access, what format it must produce, what rules it should follow, what requires escalation and how the result will be verified.
Ask whether the work can be bounded and verified, not whether lawyers regard it as intellectually sophisticated.
Where agents remain much harder to trust
At the opposite end are legal activities where the correct action depends heavily on incomplete information, tacit client preferences, competing objectives or strategic judgment.
Client counseling
Clients rarely come to lawyers because they need a technically correct description of the law. They want to know what to do. A general counsel may ask whether the company should litigate, settle, disclose, terminate, renegotiate or accept a risk. Those decisions involve law, but they also involve business objectives, personalities, reputation, risk tolerance, timing and incomplete information. The legal answer is only one input into the decision.
Negotiation
An agent may draft proposed language or compare a counterparty’s changes with a playbook, but negotiation involves much more than identifying the theoretically preferred clause. A lawyer may concede one issue because another issue matters more, recognize that pushing a provision will jeopardize the deal, deliberately leave an ambiguity unresolved or infer from opposing counsel’s behavior that a stated position is not actually firm.
That kind of judgment is difficult to encode in a static playbook.
Novel legal analysis
Agents are strongest when there is authoritative information available to retrieve and apply. Novel questions create a different problem. The relevant authorities may conflict. There may be no directly controlling precedent. The lawyer may need to reason by analogy, understand institutional behavior and predict how a regulator or court will react.
AI can assist with that analysis. That is very different from allowing it to own the conclusion.
Strategic litigation decisions
Which witness should be deposed first? Should the client seek an injunction? Should counsel make a particular argument now or preserve it? Will aggressive discovery improve leverage or antagonize the judge? These decisions depend on a developing matter rather than a closed universe of documents.
The biggest misconception: automation does not require autonomy
This may be the conceptual mistake that causes the most confusion around legal agents. People often assume there are only two possibilities: either a lawyer does the work or the AI does the work autonomously. There is an enormous middle ground.
An agent might perform 70% of a workflow, stop at a defined review point and ask a lawyer to approve the next action. It might autonomously extract information but require human approval before interpreting the commercial consequence. It might draft a redline but be prohibited from sending it to opposing counsel. It might conduct research but require citation verification before its conclusions can be incorporated into advice.
It might prepare a diligence report but highlight uncertain findings for manual review. That is still meaningful automation. In many cases, it may be the better architecture.
Human-in-the-loop should become human-at-the-right-point
The phrase human in the loop is often used as if having a lawyer somewhere in the process automatically makes an AI workflow safe. It does not. If a lawyer receives a 100-page AI-generated report and is expected to rubber-stamp it in five minutes, the presence of a human has not created meaningful oversight.
The real design question is: At which points in the workflow does human judgment materially reduce risk? That may mean inserting review before a conclusion is communicated to a client, a contractual position is sent to a counterparty, a document is filed, a high-risk issue is classified, an agent changes records, a regulatory interpretation becomes operational policy or an uncertain result determines the next step in the workflow.
The objective should not be maximum autonomy. It should be maximum useful delegation with strategically placed human control.
Professional responsibility makes this more than a product-design question
There is also a reason legal agents cannot be evaluated solely by their technical capabilities. Lawyers remain professionally responsible for the work performed with AI. ABA Formal Opinion 512 applies existing duties of competence, confidentiality, communication, supervision, candor and reasonable fees to lawyers’ use of generative AI.
Those obligations become even more important as AI moves from answering questions to taking actions. The more autonomy a system is given, the more important it becomes to determine what authority it has, what information it can access, what actions require approval, what work must be verified, what activity is logged, how errors are escalated and who remains accountable for the final work product.
An agent with access only to a closed collection of contracts and permission to create a draft internal diligence table presents one category of risk. An agent permitted to communicate directly with a client, send a redline to opposing counsel or file something with a court presents a fundamentally different category.
Agentic capability and delegated authority are not the same thing.
A system may technically be capable of taking an action without the firm deciding that it should be permitted to take it.
This is where legal engineering becomes important
The rise of legal agents also strengthens the case for legal engineering. A useful legal agent is not created simply by giving a powerful model a prompt saying, “Handle this matter.” Somebody still needs to design the workflow.
That means answering questions such as:
What are the inputs?
Which sources are authoritative?
What tools can the agent use?
Which decisions are deterministic?
Which decisions require legal judgment?
What constitutes a failure?
When should the agent stop?
When must a lawyer intervene?
How will the result be evaluated?
How will errors be detected?
What needs to be logged?
What happens when the normal workflow encounters an exception?
Those are legal-engineering questions. The more capable agents become, the more important workflow design becomes.
Agents may expose something lawyers have ignored for years
To automate legal work, firms first have to understand how that work is actually performed. And many firms do not. A process may live partly in a precedent, partly in a partner’s memory, partly in an associate’s checklist and partly in an unwritten convention that everyone on the team somehow learns over time.
An AI agent cannot reliably execute an undefined process. So the adoption of agents may force law firms to make tacit legal knowledge explicit. They will need to identify standard workflows, decision points, escalation criteria, approved precedents, review requirements, quality standards and the circumstances in which normal procedures should not apply.
In that sense, legal agents may create value even before complete automation is achieved. They force organizations to engineer their legal work.
The benchmark problem is not solved either
We should also be careful not to treat new legal agent benchmarks as definitive rankings of which system is “best.” They are an important improvement over isolated question-answer tests, but agent performance is unusually sensitive to the environment around the model.
The same foundation model can perform differently depending on the tools it receives, the instructions it is given, how documents are retrieved, how memory is managed, how many iterations it is allowed, whether intermediate work is checked and how the agent harness decides what to do next.
Harvey itself has noted that model selection, harness optimization and post-training can materially affect legal agent performance. This means legal AI evaluation increasingly needs to test systems, not merely models.
Firms should stop asking whether an agent is “accurate”
Accuracy is too broad a word. When evaluating a legal agent, firms should ask much more specific questions about the exact workflow being delegated and how the system behaves when that workflow becomes difficult.
What exact workflow has been tested?
An agent that performs well on contract extraction has not thereby proven itself capable of litigation strategy. Firms should assess systems against the actual workflows they intend to automate.
What is the end-to-end completion rate?
A benchmark showing that individual components are highly accurate can still conceal a much lower success rate for the completed work product. End-to-end performance matters.
What happens when the agent is uncertain?
Does the system flag uncertainty, ask for more information, stop and escalate, or simply continue with a confident answer?
Can important conclusions be traced to a source?
Verifiability becomes more important as the agent performs more work. Lawyers should be able to trace important conclusions back to the materials supporting them.
What actions can the agent take without approval?
Technical capability should not automatically become operational authority. Firms need explicit permissions and escalation rules.
How does the system handle exceptions?
Real legal matters rarely follow the happy path. A useful evaluation needs to test abnormal documents, missing information, conflicting instructions and other edge cases rather than only ideal examples.
A better framework: do, review, decide
One of the simplest ways to think about legal agents today is to divide legal work into three layers.
1. Work the agent can do
This is the execution layer: search, extract, organize, compare, classify, populate, draft, monitor and apply a defined playbook. These activities can increasingly be delegated to agents in appropriate circumstances.
2. Work the lawyer should review
This is the verification layer. Is the source correct? Did the agent miss anything? Does the extracted provision mean what the agent says it means? Is the draft consistent with the client’s position? Have the authorities been interpreted correctly? Does an exception apply?
The lawyer is checking the quality and completeness of the agent’s work.
3. Work the lawyer should decide
This is the judgment layer. What risk should the client accept? What should we negotiate? What should we disclose? Should we litigate? Should we settle? Which issue matters most commercially? What advice should ultimately be given?
As agents improve, the boundary between the first and second categories will move. The third category will also be affected by AI, but assistance and decision-making are not the same thing.
The future legal team may look very different
The most plausible future is therefore not a law firm in which every associate is replaced by one enormous autonomous agent. It is a legal team in which humans and specialized agents work together.
An intake agent organizes the matter.
A diligence agent reviews contracts.
A research agent investigates legal questions.
A drafting agent prepares first drafts.
A negotiation agent compares redlines against a playbook.
A knowledge agent retrieves relevant precedents and institutional knowledge.
Human lawyers would supervise the system, deal with exceptions, counsel the client and make strategic decisions. The lawyer may increasingly become the orchestrator and reviewer of legal production, rather than the person manually performing every component of it.
That is still a profound change.
It will also change what lawyers are valuable for
If agents increasingly perform retrieval, extraction, first-pass review and drafting, then the value of a lawyer cannot continue to depend primarily on being faster at those tasks than another human. The lawyer’s comparative advantage shifts toward judgment, problem framing, client understanding, risk assessment, negotiation, strategy, quality control and the ability to design and supervise systems that produce legal work.
That does not make legal expertise less important. It makes expertise more concentrated at the points where it matters. The lawyer who understands both the law and how to structure a legal workflow may become considerably more valuable than the lawyer whose primary advantage is the ability to manually execute every step.
So, are legal AI agents ready for real legal work?
Yes. But that answer needs an important qualification. They are ready for real legal work in the sense that they can already perform meaningful components — and increasingly substantial sequences — of work that lawyers previously performed manually.
They are not yet ready for unrestricted autonomous legal practice. The new long-horizon benchmarks are valuable precisely because they expose the gap between those two statements. Legal agents can perform impressive amounts of an assignment while still failing the standard required for the final work product.
That is not evidence that the technology is useless. It is evidence that the right deployment model today is not “AI replaces the lawyer.” It is that the legal workflow is redesigned around what the agent can reliably execute, what the lawyer must verify and what remains a human decision.
The first generation of legal AI asked lawyers to become better at prompting machines. The agentic era will ask law firms a harder question: how much of legal work can be turned into a system — and where, inside that system, does the lawyer still need to be indispensable?
Last updated: August 2026. Analysis and educational content only; not legal advice.