Published October 5, 2026 · by Anže Skodlar
AI Glossary for Lawyers: From Model to Hallucination
The key AI concepts a lawyer needs to evaluate legal AI tools — from large language models and tokens to hallucinations, RAG, and agents.
A lawyer doesn't need to know how to code to use artificial intelligence well. They do, however, need to understand a few basic concepts, because these concepts come up exactly where their judgment is decisive: in vendor proposals, in data processing agreements, in internal policies, and in discussions about whether a particular tool is even suitable for a legal department or law firm.
The difference between a lawyer who knows these concepts and one who doesn't is not that the former knows more. It's that they know how to ask the right question and recognize when a vendor's answer isn't really an answer.
The overview below is organized in levels: first the umbrella concepts, then how a model works with text, then typical errors, and finally the architectural solutions that address them.
Level One: What Is Artificial Intelligence?
Artificial intelligence (AI): An umbrella term for systems that perform tasks that would otherwise require human understanding (pattern recognition, classification, writing, reasoning, etc.). The term on its own says little: it covers everything from a simple spam filter to systems that draft legal texts. So when a vendor says "we use AI," that isn't information. It's a starting point for asking what kind of system, exactly.
Machine learning: An approach in which a system isn't programmed with rules but learns from data. Instead of someone writing the rule "if the contract contains this word, file it here," the system extracts the pattern itself from a large number of examples. The consequence that matters for a lawyer is that such a system has no written rules that could be read and checked.
Large language model (LLM): A model trained on enormous amounts of text that, on that basis, predicts which word is most likely to follow the text so far. Everything we know today as conversational tools (ChatGPT, Claude, Gemini, etc.) is, at its core, a large language model. The key thing to understand is what the model does: it doesn't look up an answer, it composes one. Most of what follows stems from this single property.
Generative AI: A collective name for systems that create new content (text, images, audio, code) rather than merely classifying or searching existing content. It is legally relevant mainly because generated content raises questions of authorship, labeling, and liability that don't arise with a search engine.
Level Two: How a Model Works with Text
Prompt (input instruction): Everything you enter into the tool: the question, the attached contract, instructions on what to do with it, and so on. Apart from the sources the tool has access to, this is the only thing it knows about your matter. The quality of the prompt is therefore not cosmetic. It is the input on which the output depends.
Token: The basic unit into which a model breaks text, roughly a piece of a word. Tokens measure both the amount of text a model can process at once and the cost of use. For Slovenian, this is not neutral. Models learned their tokenization rules from material heavily dominated by English, so an English word often takes up a single token, while a Slovenian word is split into several pieces, and not only when it is long. Declensions and diacritics (č, š, ž) amplify the effect. The same text in Slovenian therefore takes up noticeably more tokens than in English; exactly how many more depends on the model.
Context window: How much text a model can have "in front of it" at once: the prompt, attached documents, and the conversation so far, all together. Imagine a desk on which you can lay a limited number of sheets: what's on the desk, the model sees; what isn't, doesn't exist. Modern models can handle anywhere from a few dozen to several hundred pages of text, depending on the model. For a lawyer, this means that if you upload a case file that exceeds the window, the model hasn't "forgotten" anything. It simply never saw that part. And as a rule, it won't tell you so on its own.
Knowledge cutoff: The date up to which the data the model was trained on extends. Anything later is unknown to it: an amendment to a statute, a change in case law, a new judgment. The trap is that the model doesn't sense this boundary: asked about a law that has since been amended, it will answer confidently, based on the old version. That's why, in legal work, it is essential whether the tool also has access to current sources in addition to the model.
Level Three: Where Things Can Go Wrong
Hallucination: An answer that looks correct but doesn't correspond to reality (e.g., a fabricated judgment, a wrong case number, a quotation that isn't in the provision). This isn't an error in the sense of a malfunction. It is a direct consequence of what the model does: if the correct judgment isn't in its memory, it will compose the one that is most probable given the patterns. The model has no sense of not knowing something, so it can't tell you that it doesn't know.
Sycophancy: The model's tendency to align itself with the position the user already carries in the question. "Is it true that I can terminate the contract in this case?" and "Can I terminate the contract in this case, and what speaks against it?" often don't yield the same answer. In legal advice, this is one of the more serious pitfalls, because it precisely doubles the confirmation bias the lawyer is already exposed to. The solution is simple: also ask for the counterarguments, not just for confirmation.
Level Four: How It Gets Solved
RAG (retrieval-augmented generation): An architecture in which the system first searches a collection for relevant documents and only then composes an answer from what it found. Put simply: generation backed by a search of sources. The difference is fundamental, because instead of "what the model remembers," we get "what the system found." This doesn't eliminate hallucination entirely, but it changes the nature of the error: a lawyer will notice a misreading of a retrieved document while reading it, but not a fabricated judgment with no source.
Vector search and embeddings: A search method in which text is converted into numerical representations that capture meaning rather than words. That's why such a system will find a judgment dealing with the same legal situation even if it uses different terminology from your question, which, when searching case law, is often exactly what you need. Classic keyword search can't do this.
Traceability or grounding of an answer: The property of a system to attach every claim in its answer to a specific source that can be opened and read. For a lawyer, this is practically the only criterion that counts: no one is going to assess the inner workings of the model, but whether an individual claim is verifiable is a question that can be answered with yes or no.
Agent: An agent is a system to which, instead of a single question, you assign a goal; it then determines the steps toward it on its own, carries them out in sequence, and reports back only with the result. Given the task "compare these three contracts with our template and prepare a list of deviations," an agent would itself open each document, find the relevant clauses, compare them, and compile the list. With a conversational tool, you would guide those steps yourself, provision by provision. So with an agent, you're not trusting an individual answer, but a sequence of its decisions. In legal work, it is useful when the individual steps are visible and verifiable, not just the final result.
Putting the Concepts into Practice
These very concepts are the yardstick by which Veru, legal AI for the Slovenian market, is built. Answers are generated from retrieved legal sources, not from the model's memory. Every claim is attached to a source that can be opened and read, and currency doesn't depend on the model's knowledge cutoff. Users' inputs and documents are not used to train or improve models. Verification against the original source remains the lawyer's step.
Concepts only gain meaning once you use them. So here's our suggestion: take one question from a matter you're currently handling, ideally one where you already know the correct answer, and put it to Veru. Then open the cited sources and check whether they say what the answer claims.
A few minutes of this kind of test will tell you more about a tool than any description of its features. And afterward, you'll know how to check every other tool someone offers you in the same way.