How do we limit the scope of data sent to AI?
The AI agent does not get open access to the entire firm or all case files.
Documents, pleadings and recordings are stored in a secure environment and prepared for semantic search.
Before calling the LLM, the system searches the vector database only for fragments related to the specific question.
The model receives limited context needed to perform the task, not the firm's entire archive.
AI supports the lawyer, but does not take over responsibility for legal assessment and decisions.
Artificial intelligence can significantly speed up a lawyer's work. It can help find information in documents, summarize meeting findings, draft a letter, identify relevant parts of materials or answer a question about an ongoing case.
In a law firm, however, there is a question more important than convenience itself: how much information should an AI model receive to perform a specific task? Our answer is simple: as little as possible.
The agent should not receive entire case files just because a lawyer asked one question. First, the system determines what information is needed to answer it, and only then can selected fragments be sent to the language model.
AI itself is not the problem. Uncontrolled access to data is the problem
The discussion about artificial intelligence security often comes down to whether law firm documents can be sent to AI. That question is too broad. What matters more is: which documents, which fragments, for what purpose and to what extent should be shared?
If a lawyer asks about the client's arrangements on the payment deadline, the answer may require two paragraphs from the contract, a fragment of a meeting note and part of a call transcript. There is no reason to send the client's entire history with them, all case documents, materials from other proceedings, recordings unrelated to the question or information about other clients.
The model should receive only what it actually needs. This is what we call context control.
First, data stays in the law firm's secure environment
Pleadings, documents, recordings and other materials created and stored in Bezpieczna Kancelaria are kept in an isolated environment and encrypted.
- documents uploaded to the system
- pleadings created by the law firm
- case related materials
- recordings
- transcriptions
- content later used by AI features
Storing data is only the first element of the architecture. If the system is later to help a lawyer find an answer in hundreds of documents, it must be able to recognize which fragments are related to the question asked.
Documents are prepared for semantic search
Content stored in the law firm is vectorized. Semantic search makes it possible to find information related to the meaning of the question, even if the document uses different wording.
The information needed for this search goes to a vector database. Its role is not to replace documents, but to help the system find the parts of the materials that may be needed to perform a specific task.
The agent does not freely search the entire law firm
When a user asks the agent a question, we do not send the model the entire data set and ask it to choose the information it needs. The Bezpieczna Kancelaria system acts first.
- the lawyer asks a question
- the system analyzes what information the answer may require
- a search is performed in the appropriate data set
- the vector database identifies the most relevant fragments
- the system prepares limited context
- only this context is then sent to the language model
- the model prepares an answer for the user
The model does not receive the keys to the entire archive. The system brings it only the pages that may be needed to perform the specific task.
The user sees an intelligent agent. Underneath, a controlled system is working
A lawyer can ask how the client justified refusing to pay and, moments later, receive an answer that takes into account a document, note or conversation stored in the system. This may create the impression that the agent knows the entire case, all documents and the firm's history.
In practice, the agent may receive only a few fragments selected in advance by the system. Security therefore depends not only on which AI model was used. What happens before the model receives the question is extremely important.
We control the agent, not the other way around
The model receives a broad range of data and decides which information to use for the answer.
The system first limits the data, selects the needed fragments and only then prepares context for the model.
We do not design a solution where the model receives a full picture of the law firm and only then considers what to use from it. First, we limit the data. Only then do we use the model's capabilities.
Professional secrecy and the principle of information minimization
Lawyers have long known the rule that confidential information should not be disclosed to a person who does not need it to perform their task. A similar way of thinking can be applied to artificial intelligence.
- if a fragment of one document is enough to answer the question, there is no need to share the entire file
- if the question concerns one case, the context should not include materials about another client
- if a short finding from a transcript is enough, there is no need to share the entire recording
This does not mean that no information can ever be sent to a language model. It means that only the information needed to perform a specific task should be sent to the model.
AI should not get extra information just in case
Traditional IT systems often use the principle of least privilege. We apply a similar principle to AI. The model should not receive additional data only because it might prove useful someday.
The question before calling the model should not be: what else can we send it? It should be: how little data is enough to perform this task correctly?
Example: analyzing a client call
If the law firm has a recording of a one hour client call and the lawyer asks about the planned date for selling a property, there is no need to send the entire recording to the model. The system can find parts of the conversation about the sale and the date, then send the model only the selected context.
Example: question about a pleading
A lawyer can ask what arguments about limitation periods have appeared so far in a given case. The system can find the most related fragments of pleadings, notes or other materials connected with the indicated case. Only those fragments are then used to prepare context for the model.
AI security starts before AI
It is easy to focus the discussion on the language model itself: who created it, where it runs, what terms it has and how long it stores data. These are important questions, but the security of an AI system in a law firm starts earlier.
- where the documents are located
- how they are protected
- how the system searches for information
- who can trigger a given action
- what scope of data is searched
- which fragment becomes context
- when information may be passed on
If the entire archive is sent to the model from the start, later restrictions mean little. That is why in Bezpieczna Kancelaria we first limit the context. Only then do we ask the model.
AI should help lawyers without taking over their responsibility
Even the best secured agent remains a tool. It can help quickly find information, organize facts, summarize materials and prepare a working answer. However, it should not replace a lawyer's professional assessment.
- checking the source
- reading the full document
- assessing the procedural context
- verifying the legal status
- the decision of the person handling the case
Lawyers should know what happens to their data
When choosing an AI system for a law firm, it is worth asking not only whether it has artificial intelligence features, but above all how it limits data access and prepares context.
- Does the agent get access to entire case files?
- Does the system limit context before sending the query?
- Can data from different clients be accidentally combined?
- Does the search take place before the LLM is called?
- Can you control the scope of information used for the answer?
- How are documents and recordings protected from use by AI?
- Is the model a tool of the system, or does the system hand over control of data access to it?
The agent should know only what it needs to know
Artificial intelligence can be extremely useful in a lawyer's work. But this does not require giving it all the knowledge of the law firm.
In BezpiecznaKancelaria.pl, documents, pleadings and recordings remain in a secure environment. Their content can be prepared for semantic search. When a user asks a question, the system first finds the most relevant information, prepares limited context and only then sends it to the language model.
In a law firm, the question should not be: can AI read all my documents? It should be: which information does AI really need to see to perform this one task?

