Writing

Classical ML or an LLM? The decision before touching code

In short: if the input is numbers or a signal and you have labelled examples, I start with a classical model. If the input is language and the job is to read or write it, I start with an LLM. Most of…

published
read time
5 min
words
908
lang
en
filed under
Engineering

In short: if the input is numbers or a signal and you have labelled examples, I start with a classical model. If the input is language and the job is to read or write it, I start with an LLM. Most of the real work is being honest about which of those two problems you actually have.

Every few weeks someone asks me whether they should "use AI" for a problem, and lately that means an LLM. Sometimes the answer is yes. Often the answer is a gradient boosted tree that trains in a minute and nobody finds exciting. The choice should happen before anyone opens an editor, because it decides the data you collect, the budget, the evaluation and what you'll be able to explain later.

A note on words. By classical I mean a model trained on your own labelled data for one task. That covers logistic regression and random forests, and for this post also small task-specific neural networks. By LLM I mean a general language model you prompt, with or without retrieval, and with or without fine-tuning.

Start with three questions

  1. What goes in? A row of numbers, an audio clip, an image, or text that a person wrote.
  2. What comes out? A label, a number, or text that a person will read.
  3. What does a wrong answer cost? And who has to explain it when it happens.

Most of the time those three answers make the decision for you. In health, where I've spent most of my work, the third question carries the most weight. Being wrong has a cost, and someone will ask why the model said what it said.

Two defaults

What each one gives you

Classical

  • Needs labelled data
  • Fast and cheap per prediction
  • Same input, same output
  • Feature-level explanations with SHAP or LIME
  • Runs anywhere, including on-site
  • Weak on free text

LLM

  • Works from a prompt and a few examples
  • Slower, paid per call
  • Output can vary between runs
  • Explains itself in words, which isn't the same as a reason
  • Data often leaves your servers
  • Strong on language, weak on arithmetic

Neither column is better. They're good at different shapes of problem.

The decision rules I use

QuestionLeans classicalLeans LLM
What is the input?Numbers, signals, imagesFree text, documents, conversation
Do you have labels?Hundreds or more, and trustedFew or none
Must the answer be repeatable?Yes, audited or regulatedSome variation is acceptable
What explanation is needed?Which inputs drove the decisionA quoted source or a readable summary
Volume and latencyMany calls, millisecondsModest volume, seconds are fine
Can the data leave your servers?NoYes, or you can host a model
How will you measure it?Held-out set and one metricA rubric and human review

When the rows disagree, the first two usually win. A model that matches the shape of your input and the labels you actually have will beat a better-sounding model fighting against both.

How the decision plays out

Voice screening: classical side

Take a typical voice-screening task, where a model listens to a recording and flags a condition. It is audio in and a label out. There are labelled recordings, the decision has to be repeatable, and a clinician will want to know what the model heard. That's a trained model on acoustic features or spectrograms. A language model has nothing to add to the core prediction.

Clinical decision support: classical side

At a hospital network I built decision support with federated learning, where the data stayed at each site. Tabular clinical features, a label, and a hard rule that patient data doesn't travel. Sending records to a hosted LLM wasn't an option, and feature-level explanations, the kind SHAP gives you on a tree model, are something a clinician can actually read and argue with.

A companion that talks: LLM side

When I built a voice companion for patients with dementia, the whole product was conversation. Nobody has a labelled dataset of good replies to every sentence an older adult might say. That was an LLM from the first day, and the engineering went into speech recognition, voice and privacy around it.

The hybrid that keeps winning

The most useful pattern sits in the middle. Use an LLM to turn messy text into structured fields, then hand those fields to something simple and testable. I did a version of this for literature reviews: the model reads a title and abstract and answers include or exclude against written criteria, and the result is a plain list you can check against a hand-labelled sample. The LLM does the reading. The bookkeeping and the measurement stay in plain code.

Signs you picked wrong

  • You're writing longer and longer prompts to make an LLM produce a number. That's a regression problem. Train a regressor.
  • You're hand-crafting dozens of text features for a classical model and it still misses the point of the sentence. Let a language model read it.
  • You can't say how you'd evaluate it. That's not a model choice problem yet. Go back to the three questions.

Next time a project lands on your desk, write the three answers on one line before you write any code: what goes in, what comes out, what wrong costs. Then go through the table row by row. If you still can't decide, build the classical baseline first. It takes an afternoon, and it gives the LLM a real number to beat.

related

Keep reading