Rules vs transformers for extracting metadata from legal documents
In short: on the same Spanish-language legal documents, the transformer path was more accurate and slower. The rules were fast, cheap and easy to explain. I kept both: rules by default, the…
- published
- read time
- 5 min
- words
- 968
- lang
- en
- filed under
- Engineering
In short: on the same Spanish-language legal documents, the transformer path was more accurate and slower. The rules were fast, cheap and easy to explain. I kept both: rules by default, the transformer on request, and an endpoint that runs the two side by side so I can see where they disagree.
The job: three fields from a pile of documents
The task sounds small. Take a legal document from one of several Spanish-speaking jurisdictions and pull out three things:
- Document type, one of a handful of categories.
- Issuing body, both as written and in a canonical form, so that the long formal name of one chamber of a high court maps to one clean label.
- Document date, as day, month and year.
Each field comes back with a confidence score. The documents range from a single page to long filings. And the dates are often not dates. They are sentences. A ruling can open with a city name and then "catorce de diciembre del año dos mil once", which is the fourteenth of December 2011 written out in words. A regex for dd/mm/yyyy finds nothing there.
I built this as a small Python service with one extractor per field. Then I built it twice.
Two ways to get the same three fields
The first version is rules plus a statistical spaCy model for Spanish. The second uses transformer models. Same input, same output schema, different insides.
Side by side
Rules and spaCy
- One config file per jurisdiction, with patterns and weights
- A helper that turns Spanish number words into digits
- A dictionary that maps entity names to canonical labels
- Fast, runs anywhere, every decision traceable to a line of config
- Misses phrasing nobody wrote a rule for
Transformer
- Reads the field from context, not from a pattern
- Handles phrasing the rules never saw
- Slower, and heavier to run
- Harder to say why it picked a value
- Adding a jurisdiction is not a config change
Why I like the config files
The thing I like most about the rules version is the config. Adding a jurisdiction means adding one config file, not code. A lawyer who knows how one country's courts write dates can read that file and tell me what is wrong with it. Nobody can do that with model weights.
Running both on the same text
Opinions about rules and models are cheap. What I wanted was a list of documents where the two paths gave different answers, because that list is where all the learning is. The service has a compare endpoint for this. Here is the same idea as a small client, so you can see what is being compared:
def key(result):
return {
"type": result["type"]["value"],
"body": result["body"]["canonical"],
"date": result["date"]["value"],
}
def disagreements(text, jurisdiction, rules_extract, model_extract):
a = key(rules_extract(text, jurisdiction))
b = key(model_extract(text, jurisdiction))
return {k: (a[k], b[k]) for k in a if a[k] != b[k]}
Run that over a folder, write the non-empty results to a file, and read them. Not the counts. The actual documents. Time the two calls while you are at it. The gap between them is the whole speed story.
The honest result
The transformer was more accurate. It was also slower. That is the result, and I am not going to dress it up with a percentage I did not measure on a labelled set. What I can give you is where each one was better.
| Approach | Accuracy | Speed | When it goes wrong | What you can inspect |
|---|---|---|---|---|
| Rules and spaCy | Good on clean, typical documents | Fast | Phrasing no rule covers | The pattern that fired and its weight |
| Transformer | Higher, especially on odd phrasing | Slower | Confident on the wrong span | A score, not a reason |
The pattern behind that table is not mysterious once you look at what each field asks for:
- Dates suit rules. Once a number-word converter exists, a written-out date is a parsing problem, not a reading problem. Legal Spanish is formal, and formal text repeats itself.
- Issuing bodies are hard for both. Finding the raw string is the easy half. Mapping it to the right canonical name is a dictionary problem, and a model does not fill in a missing dictionary entry.
- Long documents favour rules. The date and the issuing body usually sit near the top. Rules can look at the first page and stop. A model that reads a long filing end to end pays for pages it does not need.
So the default endpoint stays rule-based. The transformer is there when someone asks for it, and the compare endpoint is there for me.
How I would set this up again
If you are extracting structured fields from legal text, here is the order I would do it in:
- Write the output schema first, with a confidence score on every field. Both methods have to fill the same shape, or you can never compare them.
- Build the rules version. It is your baseline and your documentation of how the documents are written.
- Put the jurisdiction-specific parts in config, not in code. You will add a jurisdiction sooner than you think.
- Add the model path behind a separate endpoint, not as a replacement.
- Log every disagreement between the two and read them by hand. Each one is either a missing rule or a model mistake, and both are worth knowing.
- Only then decide what runs by default. A good next step is to send just the low-confidence rule results to the model, so you pay the slow path only where the fast one is unsure.
Start with step five even if you skip the rest. A list of disagreements will teach you more about your documents than any benchmark.
related