Choose Per Job, Not Per Person
"Local or cloud?" is too large a question. A team may summarise sensitive documents, classify routine records, draft occasional proposals, and ask a demanding model to solve an unusual problem. Those jobs have different boundaries.
Decide where one workload should run. Local wins when data control, steady volume, offline use, or direct version control dominates. Cloud wins when model capability, burst capacity, shared access, or low infrastructure burden matters more.
This is the practical extension of treating AI as infrastructure: define the job, test the route, and keep the decision record.
Choose the Kitchen for Each Service
Think of the decision as choosing a kitchen. Predictable preparation can stay in-house. Difficult work or a sudden banquet may need specialist equipment and extra capacity elsewhere.
The useful rule is not "always in-house" or "always outsource". It is which ingredients may leave, what quality the result must reach, and what the full service costs.
The analogy has limits: data can be copied and AI output can vary. But the operating lesson holds. Pick a location per job and document the boundary.
Steps
Guide
Define the Job Before the Platform
Name one job, its owner, its input, and the output it must produce. Record whether demand is constant, scheduled, occasional, or bursty.
Do not start by comparing model names. Models, apps, and harnesses are different layers, and the deployment decision sits across them.
- What must the output let someone do?
- How often does the job run?
- How quickly must it finish?
- What happens if the route is unavailable?
A narrow job produces a useful decision. "Run our AI locally" does not.
Draw the Data Boundary First
Classify the source data before testing any remote service. Write what may leave the device or private network, what must be redacted, and where outputs may be stored.
A local model can keep prompts on controlled hardware. Ollama's official FAQ says local prompts and data are not visible to Ollama, and explains how to disable its optional cloud features.
But a local-looking interface may still call web search, plug-ins, telemetry, or a hosted model. Verify the whole route, not only the model label. Include retrieval sources, logs, backups, extensions, and output storage.
Cloud is not one retention policy either. OpenAI's API data controls show how training use, abuse-monitoring logs, and application state can differ by endpoint and configuration. Read the current terms for the exact service you plan to call.
Use Local When Control Earns Its Cost
Local execution is the clearest fit when approved data must not leave controlled hardware. It also fits work that must continue without reliable internet access.
Steady, high-volume work can favour local operation because another run does not create another provider token bill. That does not make it free. Electricity, hardware depreciation, monitoring, updates, backups, and staff time still count.
Local control also lets a team hold a model artefact and runtime version unchanged. This can improve repeatability when next month's output must be checked against this month's setup.
Control is not determinism. Keep the inputs, settings, and representative outputs. Even a pinned setup needs evaluation when something changes.
Use Cloud When Capability or Demand Moves
Cloud services provide models and compute that may not fit on a laptop or office workstation. Quantisation reduces model memory requirements, but it changes the exact model setup being tested. A smaller local route is useful only if it still clears the business quality threshold.
Cloud also suits work that spikes and stops. AWS guidance identifies transient workloads and workloads needing newer hardware as candidates for cloud elasticity.
Shared work is easier to centralise behind one managed endpoint. Several people, devices, or automated systems can use the same access controls and service policy.
The provider runs the underlying infrastructure, but the customer still owns access control, integration, evaluation, and data policy. Managed infrastructure does not remove governance.
Count the Costs That Appear in Different Places
Local and cloud expose cost differently, so compare both over the same period and workload volume.
For local execution, count hardware, expected useful life, electricity, support, monitoring, backup, maintenance, replacement risk, and staff hours. Utilities, support, facilities, and staff time belong in the total cost of ownership even when nobody sends an invoice for them.
For cloud execution, count inference, tools, storage, regional or retention options, integration work, operational review, and network transfer. Google Cloud's storage pricing is one example of data transfer appearing as a separate cost category.
Do not compare a token rate with a purchase invoice. There is no universal break-even point. Existing hardware, workload volume, electricity, staffing, model choice, and provider terms change the result.
Run the Same Work Both Ways
Use representative inputs, the same instructions, the same output format, and one review rubric. If the local route uses a quantised model, test that exact version. If the cloud route needs retrieval or tools, test the complete route.
Keep both outputs. Review correctness, usefulness, consistency, elapsed time, failures, and the human repair needed. Record the cost units consumed and the operator time.
A pinned cloud snapshot can improve repeatability too. OpenAI recommends pinned model versions and evaluations when consistent prompting behaviour matters, while noting that outputs remain variable.
The winning route is the one that clears the business threshold with an acceptable boundary and total cost. It does not need to win every benchmark. This is the same practical reason to use different AI apps for different jobs.
Write the Hybrid Crossing Rule
Many teams will keep sensitive or high-volume routine work local, then use the cloud for difficult reasoning or burst capacity. That is not a complete architecture until the crossing rule is written.
Name the data class, allowed destination, required transformation, owner, review point, and fallback. A restricted document might remain local while a human-approved redacted question goes to an approved cloud endpoint.
Route manually first. Automate only after the rule survives real use. A weak automatic classifier can send the wrong material across the boundary.
State what happens during failure. A cloud outage may pause a restricted job instead of switching it to an unapproved service. A local failure may send a public-data task to the cloud while restricted work waits.
Use the Browser When the Browser Is Enough
Local does not always mean a model server. Sometimes the browser can complete the whole job.
Most N-G.BE tools process image, text, and data work in the browser: conversion, counting, ranking, password generation, and colour extraction. None of those needs a model. They are local because the job is small enough to be.
The Built by meBackground Remover is the one that runs an AI model locally. It is not my model: it uses the open IMG.LY browser-side package, which downloads its runtime and ONNX model on first use, then cuts the image out on the visitor's own device. That is the trade in miniature: a one-off download and the visitor's hardware, in exchange for images that never reach a server. A specialised cloud remover may still cut hair or glass better.
WebLLM shows that model inference can also run inside supported browsers through WebGPU. Whether it should depends on model size, browser support, performance, and the quality threshold.
The browser example is not a universal privacy guarantee or a claim that demanding reasoning belongs on-device. It shows the decision method at its simplest: the workload is narrow, the available hardware is sufficient, and avoiding a server round trip creates a clear benefit.
Be Aware
The team chooses local or cloud for every workload.
Split the decision by job. Keep one short routing table that can send different work to different approved routes.
The cost comparison uses an API rate against a hardware price.
Use the same period and workload volume, then include utilities, depreciation, transfer, maintenance, and staff time.
Local is treated as automatically private.
Check the full network path and disable or remove cloud models, search, telemetry, and plug-ins that breach the written boundary.
The two routes are tested on different work.
Freeze the inputs, instructions, output format, and review rubric before comparing quality or cost.
Hybrid routing exists only in people's heads.
Record the data class, approved destination, transformation, owner, fallback, and review point before automating it.
AI Workload Decision Record
Copy / paste
Evaluate one AI workload for local, cloud, or hybrid execution.
#### Workload
* Name:
* Owner:
* Input:
* Required output:
* Frequency: constant / scheduled / occasional / bursty
* Maximum acceptable delay:
* Failure impact:
#### Data Boundary
* Data classification:
* Data that must remain local:
* Data allowed to cross:
* Required redaction or transformation:
* Approved destination and region:
* Retention and logging policy checked on:
#### Matched Test
* Representative inputs:
* Shared instructions:
* Required output format:
* Quality threshold:
* Local route and exact model version:
* Cloud route and exact model version:
* Retained output locations:
* Human reviewer:
#### Full Cost
* Comparison period and volume:
* Local hardware, electricity, maintenance, backup, and staff time:
* Cloud inference, tools, storage, transfer, options, and staff time:
#### Decision
* Approved route: local / cloud / hybrid
* Reason:
* Hybrid crossing rule:
* Fallback:
* Decision owner:
* Next review date:
Do not choose a route until the data boundary and quality threshold are clear. Do not assume local is automatically private or cloud has one universal retention policy. Use current terms for the exact services being tested.



