INSIGHTS
Governing the Router: Local Models, Frontier Access, and the New Control Point
Silvan Schriber · June 2026
The next AI decision your bank makes won't be about which model to buy — it will be about where data is allowed to leave the building. As AI splits into two layers — small models running locally on staff hardware, and frontier models locked inside regulated cloud data centres — Apple, NVIDIA/Microsoft, and Perplexity are racing to keep inference on-device, while Anthropic is doubling down on the opposite bet: governed, auditable frontier access via the hyperscalers. The dividing line carries real regulatory weight — FINMA Circular 2018/3 outsourcing obligations and revFADP/GDPR transfer rules apply the moment data crosses to the cloud, and disappear the moment it doesn't. The winning architecture isn't a choice between local and frontier — it's a tiered system where local is the default, frontier is the deliberate exception, and the router that decides which is which becomes the institution's new control point, with every handoff logged and auditable.
The terms, ordered cleanly
The vocabulary is used loosely in the market. The terms are not synonyms and they are not mutually exclusive – they describe different axes. Separating those axes is the precondition for any sensible risk discussion.
Axis A – what the model is (capability and generality)
-
Foundation model: a large model trained on broad data via self-supervision, adaptable to many downstream tasks. The EU AI Act's near-equivalent legal term is General-Purpose AI (GPAI) model. This is a category, not a size – it spans the very small to the very large.
-
Frontier model: the most capable subset of foundation models – the largest, most compute-intensive systems at the leading edge (e.g. the flagship models from OpenAI, Anthropic, Google). “Frontier” is a relative, moving label; today's frontier is next year's baseline.
So: every frontier model is a foundation model; most foundation models are not frontier models. Frontier is the sharp tip of the foundation-model category.
Axis B – where the model runs (location of inference).
-
Cloud / server model: inference runs in a remote data centre. Frontier models live here by necessity – they are too large for consumer hardware.
-
Local model: inference runs on hardware the organisation or user controls – a workstation, an on-premise server, a laptop. Data does not leave that perimeter.
-
On-device model: the narrowest case of local – inference runs on the end-user's own personal device (phone, laptop), typically a small model (~1–4bn parameters) embedded in the operating system. Apple's on-device model is ~3bn parameters.
The combinations that matter:

Source: vendor documentation (Apple ML Research, NVIDIA, Microsoft Build 2026); EU AI Act, Art. 51–55. Figures as of 06.2026.
Why the distinction is a risk distinction
The location axis maps almost directly onto the regulatory and data-protection axis. On-device inference means no third-party data processor and no cross-border transfer – the data never leaves. Cloud frontier inference means the opposite: a processor relationship, an outsourcing arrangement under FINMA Circular 2018/3, and – where the data is personal – a transfer assessed under the revFADP and, for EU data subjects, the GDPR. The model-capability axis maps onto a different regime: the EU AI Act's GPAI and systemic-risk obligations, which bind the provider of the model, not the deploying bank – but whose documentation the bank needs in its own audit trail.
The market, by layer
The supply chain has four layers. Value, and concentration risk, currently sit at the top and the bottom; the middle is contested.

Source: market structure synthesis; vendor disclosures, 2025–06.2026.
The silicon layer (NVIDIA) and the frontier labs capture most of the economics today. The OS houses (Apple, Microsoft) own the distribution and the user's data. Each layer is now pushing into the others' territory – and the cleanest expression of that is the on-device / local-model push, which lets the OS and silicon layers offer AI without paying the frontier labs per token, while keeping data on hardware they control.
Four strategies, four motives
The same underlying shift – capable small models running locally – is being pursued by four players for four different reasons. We should read these as moves on a board, not as feature announcements.
Apple – on-device foundation model as a platform
At WWDC 2025 Apple opened its Foundation Models framework, giving any third-party developer direct access to the ~3bn-parameter on-device model behind Apple Intelligence, through a few lines of Swift, free of inference cost and offline. It runs on Apple silicon (CPU/GPU/Neural Engine), with a Private Cloud Compute tier for heavier tasks. The model is deliberately not a general-knowledge chatbot – it is tuned for summarisation, extraction, structured output and tool-calling.
Motive: privacy as a product differentiator and a moat. Apple monetises hardware, not tokens. By making on-device inference free and default, it raises the cost of leaving its ecosystem and removes the dependency on external frontier labs for everyday features.
FSI relevance: any iOS/macOS app a bank builds or buys can now perform on-device text processing on confidential material with no data egress – a materially different data-protection posture from a cloud API.
NVIDIA + Microsoft – pushing AI compute onto Windows PCs and workstations
Two adjacent moves, best read together. NVIDIA has spent the past year optimising local-LLM execution on its RTX GPUs – TensorRT acceleration, Nemotron open models, and tight support for the local-inference apps (Ollama, LM Studio, AnythingLLM) that staff already download. At Microsoft Build 2026 (02.06.2026), Microsoft positioned Windows itself as an AI runtime, shipping Windows ML with TensorRT for RTX and surfacing NVIDIA RTX-class hardware (“RTX Spark”, ~128GB unified memory) able to run large models – in some cases frontier-scale – locally on a workstation.
Motive: for NVIDIA, expand demand beyond the data centre into every workstation and PC – a second, broader silicon market. For Microsoft, make Windows the default home for “agentic computing” before Apple or Google define that category, and reduce its own dependence on cloud inference economics.
FSI relevance: a bank can now run a capable model entirely inside its own four walls on commodity workstation hardware. This is the on-premise option re-emerging – full data control, no per-token cost, no outsourcing relationship – at the price of capability and operational burden.
Perplexity – hybrid routing as the product
Perplexity has made the routing layer itself the product. Its “Personal Computer” agent (Mac, 04.2026; Windows rollout in progress) and the hybrid local-server inference orchestrator announced at Computex 2026 (02.06.2026, due in Perplexity Computer from 07.2026) run a compact local model that classifies each sub-task by sensitivity and complexity, then decides – in real time, mid-task – whether it executes on-device or is sent to a cloud frontier model. The stated split: a private financial document stays local; an anonymised summary of it may go to the cloud for deeper reasoning. The orchestrator is chip-agnostic (demonstrated on Intel Core Ultra and NVIDIA RTX hardware).
Motive: resolve the three-way tension between capability, privacy and cost without asking the user to choose – and, not incidentally, push inference onto the user's hardware to cut Perplexity's own server bill.
FSI relevance: this is the architecture a bank will most likely meet in practice, and the one that most needs scrutiny. The router's classification logic becomes a control of supervisory significance: if it mis-classifies a confidential item as “safe to send,” the data has already left before any human reviews it. “Anonymised summary to the cloud” is a data-protection claim that must be independently verifiable, not taken on trust.
Anthropic – frontier capability as a governed cloud service
Anthropic is the deliberate inverse of the three strategies above. It builds only the frontier layer and has no device, silicon or consumer-platform play. There is no on-device Claude and no ambition to ship one; the model is too large to run on a phone or a standard workstation, and that is by design. Distribution runs through the three hyperscalers – Claude is the only frontier model available across AWS (Bedrock), Google Cloud (Vertex) and Microsoft – plus a direct API and the Claude Enterprise tier.
Capability is segmented by governance, not geography: a public model line for general use and a tightly gated tier (Project Glasswing) reserved for a small number of vetted critical-infrastructure partners, with no public API.
Motive: compete purely on capability, safety and trust rather than on owning the endpoint. Where Apple monetises hardware and NVIDIA monetises silicon, Anthropic monetises governed access to the most capable models. Its differentiator against the device players is precisely what they cannot offer at the frontier – leading-edge reasoning and large context – and its differentiator against other labs is the surrounding control regime: published safety cards, tiered access, and enterprise data terms positioned for regulated buyers.
FSI Relevance: Anthropic sits squarely on the side of the local–cloud line where the regulatory perimeter is most active, and its enterprise posture is built to make that side defensible:
-
Data not used for training by default on the commercial and Enterprise tiers (a contractual default, not an opt-out toggle as on the consumer tiers – a distinction the bank must police internally to avoid “shadow AI” on consumer plans).
-
Short retention and a Zero-Data-Retention option for qualifying enterprise API customers; standard API log retention reduced to 7 days (as of 09.2025), with ZDR available by addendum.
-
Deployment inside the bank’s own cloud tenancy via Bedrock or Vertex, keeping data in the institution’s VPC with regional residency controls – the practical answer to FINMA Circ. 2018/3 and revFADP transfer questions, though it does not eliminate them.
The strategies at a glance

Source: Apple ML Research & Apple Newsroom (WWDC 2025); NVIDIA developer blog (2025); Microsoft Build 2026 (02.06.2026); Perplexity / Computex 2026 (02.06.2026, VentureBeat, MarkTechPost); Anthropic (Claude Partner Network, Enterprise product pages, data-retention disclosures 09.2025–06.2026). As of 19 June 2026. “Frontier-scale local” capability claims are vendor-stated and not independently verified.
How local and frontier models complement each other
The two layers are not competitors converging on one winner. They are settling into a division of labour defined by three trade-offs that pull in opposite directions:
-
Capability favours the frontier model (cloud).
-
Privacy and data control favour the local / on-device model.
-
Cost and latency favour the local model for routine work – frontier compute is wasted on tasks a small model handles.
No single model optimises all three. The resolution is a tiered architecture in which the local model is the default and the frontier model is the exception, invoked only when a task genuinely needs leading-edge reasoning or a large context.
The architecture is the control. Once a bank accepts a tiered model, the local layer becomes a containment boundary and the router becomes the gate in that boundary. Done well, this is a stronger posture than either pure-cloud (everything leaves) or pure-local (capability ceiling): the institution keeps confidential material on controlled hardware and reaches for frontier capability deliberately, with a logged, auditable decision each time data crosses the line.