Skip to content
All articles
Automation & AI

AI in the cloud or on your own hardware: the decision is made at four points

Andrey Gershengoren · · 7 min

"Our data must not go to the cloud." A large share of conversations about automation with language models begins with this sentence, and it is rarely a decision. It is a worry looking for a decision. Whether a model runs with a provider or on a server in your own building can be settled at four points, and at each of them the answer is measurable. For many mid-sized companies it turns out differently than the worry suggests.

Local AI or cloud model: what we are talking about

Local means: an open model such as Llama, Mistral or Qwen runs on your own hardware, usually a server with one or more GPUs, operated by someone in your company or at a hoster under your control. The data does not leave the machine, the choice of model is yours, and so is the operation.

Cloud means: a model is called through a provider's interface. You pay per processed token, choose the region where processing takes place, and govern the handling of the data through a data processing agreement. The large providers now offer processing in the EU and commit contractually to not using inputs for training.

Both are tools. Which one fits is decided by four questions about your company.

Question 1: May the data leave the building?

Not "do we want to" but "are we allowed to". The answer is in your contracts with clients, in industry-specific rules, in agreements with the works council and in your own privacy policy. Anyone processing health data, client files or data under a non-disclosure agreement often finds a clear line there.

If the answer is no, the conversation ends at this point: local, with no discussion of cost or quality. If it is "yes, under conditions", those conditions can usually be met in the cloud: processing in the EU, a data processing agreement, no training on your data, and above all anonymisation of the sensitive fields before a text leaves the building. Names, addresses and customer numbers can be replaced with placeholders in your own backend; the model still classifies the case correctly. This step belongs to every pipeline I build, regardless of where the model runs.

Question 2: What quality does the task demand?

Not every task needs the strongest model available. Sorting emails by type and urgency, pulling invoice number and amount from a PDF, mapping a text onto a schema: open models in the size class that runs on a single GPU do this reliably enough, provided the pipeline has a review step for uncertain cases.

The picture changes when the model has to read long documents with nuance, reason in several steps, or write in rare languages and specialist fields. There the large cloud models still lead according to public comparisons, and the gap is noticeable to the user. I am relying here on published comparisons and on reports from others, not on my own measurements with local models; whoever makes the decision should test it on their own task with a set of real cases before ordering hardware.

Question 3: How much volume is there?

The cloud costs per processed token, so in proportion to use. Your own hardware costs once at purchase or monthly on rental, plus electricity and the time of the person who runs it. These are two different cost curves: one starts at zero and rises with volume, the other starts high and stays flat.

A simple rule follows. With small or irregular volume the cloud is cheaper, and clearly so, because the hardware would sit idle most of the time. With high, steady volume your own machine can become cheaper. Where exactly that point lies depends on the model, the length of the texts and the provider's price, and it keeps shifting because prices fall on both sides. That is why I calculate it in the project with real numbers rather than on a slide; any figure I quoted here would be wrong within six months.

Question 4: Who runs the hardware?

This question is overlooked most often and decides most often. A local model is a server with a GPU, an inference service, drivers, updates, monitoring and a backup plan. Models get superseded, drivers break, a GPU fails on a Friday evening. Someone has to own it, and not as a side task.

A company with its own IT that runs servers anyway can take on that responsibility. A company without an IT department cannot, and an external provider who looks after the machine is then once again a third party with access to all the data, only with less contractual protection than a cloud provider. Whoever answered question 1 with "no" still has to solve this one; whoever answered "yes, under conditions" has here the strongest reason to stay with the cloud.

The price, on both sides

The cloud has its price, and it is not only the one on the invoice. You depend on a provider: on its prices, on its decision to retire a model, on its availability. Every call goes over the network and takes time, which shows in interactive applications. And every transfer of data has to be justified and documented, even when it is permitted. My own pipelines run on cloud models, and I know these costs from operating them.

Your own hardware has a price of its own. You pay before the first case is processed, for a machine that is too small or too large if sized wrongly. You tie a person to its operation. The model you install today will not be the best in its class a year from now, and switching is work. And you carry the security of the machine yourself, a machine on which all the data that used to be spread across many systems is now concentrated. A poorly maintained server in-house is not a smaller risk than a well-maintained service at a provider, only a different one.

Limits

There are cases in which I recommend your own hardware despite all of this: when the data may not leave your control, the volume is high and steady, and an IT team is ready to take over operation. When all three conditions hold, local is the right choice.

In between lies a path that is often overlooked: a small local model does the pre-sorting and replaces sensitive details with placeholders, and only the cleaned text goes to a cloud model for the demanding steps. That is an architecture decision, and it requires both operating models at once. It pays off where question 1 is answered strictly and question 2 still calls for the large model.

Closing

May the data leave the building, what quality does the task demand, how much volume is there, who runs the hardware. I settle these four questions in the intro call before any talk of models, and the answers are then written down in the effort estimate.

LLMdata protectioninfrastructureautomationdecision
Contact

First conversation: 30 minutes, free of charge, no presentation.

You describe the situation, I tell you whether and how I can help. No slides, no sales pitch.