What Should Telcos Really Do About AI?

What Should Telcos Really Do About AI?

The industry is automating itself and buying GPUs. The larger prize is to own the control loop for machines that cannot afford to fail.

The telecom industry has an AI problem.

Not a shortage-of-AI problem. A strategy problem.

Every operator now has copilots, chatbots, network-optimisation pilots, hyperscaler partnerships and an “AI-native” slide in its investor deck.

But the uncomfortable question remains:

What advantage is the telco actually building?

Using a language model to summarise a trouble ticket is useful. It is not a moat.

A better chatbot can reduce customer-service costs. It does not change telecom’s position in the value chain.

Renting GPUs can generate revenue. But hyperscalers have deeper developer ecosystems, greater purchasing power and more ways to keep those GPUs busy.

Cloud should have taught the industry an important lesson: owning the route to a workload is not the same as owning the workload.

AI could repeat that mistake -this time at the edge.

The real opportunity is bigger, but also narrower.

Telcos should stop trying to become generic AI companies. They should become the trusted control plane and distributed inference fabric for the physical economy.

In plain English: decide where intelligence runs, give it the network treatment it needs, limit what it is allowed to do, and prove that the real-world action was completed.

That is a very different business from selling connectivity or compute.

Four bets are taking shape. They are not the same business.

The global operator market has broadly split into four camps.

Telstra and Telefónica are showing that AI can improve the economics of the current business. Deutsche Telekom and Verizon are moving towards closed-loop network operations. SoftBank, Singtel and Deutsche Telekom are investing in AI-RAN, edge orchestration or large-scale compute. Rakuten, Singtel, Orange and others are trying to turn automation, orchestration and trust into platforms that can be sold externally.

These approaches are often presented as different parts of one AI transformation.

They should not be.

Efficiency is an OPEX programme. It should be judged on payback and realised savings.

Autonomy is an operating-model change. It should be judged on reliability, safe resolution and the number of incidents it creates, not just the number it closes.

Infrastructure is a capital-allocation decision. It should be judged on utilisation, power, return on capital and contracted demand.

Platform is a product business. It should be judged on recurring revenue, gross margin and whether customers can buy it without a six-month integration project.

Combining all four into one “AI progress” scorecard makes weak investments look strategic.

The industry is confusing activity with position.

Most telcos are either using AI to make today’s operator more efficient, or supplying infrastructure to somebody else’s AI system.

The missing business is a repeatable product for running intelligence safely in the physical world.

What the standards are really telling us

The significance of 3GPP’s AI work is easy to lose in specification numbers.

The signal is simpler.

The model is becoming a managed network resource.

TS 28.105 gives operators a framework for managing model training, deployment, inference and lifecycle. That means a model should have a visible version, operating history, deployment location, performance envelope and rollback path. It should no longer disappear inside a proprietary vendor feature.

The radio is becoming partly learned rather than entirely programmed.

Learning-based systems are entering areas such as mobility, scheduling, beam management, channel estimation and energy control. A poor model is no longer limited to giving an engineer a bad recommendation. It can affect network behaviour directly.

And the network is becoming part of the computer.

3GPP has assigned Release 20 to 6G studies and Release 21 to normative 6G work. The ITU’s IMT-2030 framework includes both AI and Communication and Integrated Sensing and Communication as usage scenarios. O-RAN Release 5 extends AI-model workflows across the non-real-time and near-real-time RAN Intelligent Controllers.

Until now, a network mainly answered two questions:

Where should these bits go?

And how quickly can we move them?

The next architecture will increasingly have to answer four more:

Where should this inference run?

What application or environmental context should affect network behaviour?

What action is the system permitted to take?

And did that action happen within its operational deadline?

That is the real strategic meaning of AI-native 6G. It is not simply a faster network running more models. It is the convergence of connectivity, compute, sensing and control.

Three traps sit between today’s pilots and that future

Prediction is not permission

A model may correctly predict congestion. That does not mean it should receive permission to alter a national network.

A maintenance agent may identify the likely cause of a fault. That does not mean it should hold unrestricted credentials across the RAN, transport, core and cloud.

Operators need a hard separation between intelligence and authority.

A policy layer and not the model must decide what can be changed, over which sites, at what confidence, with what maximum impact, and under which canary and rollback conditions.

The strategic control point is therefore not the foundation model. It is the system that turns evidence into authorised action and can later reconstruct exactly why that action occurred.

A GPU inside an exchange is not a business model

AI-RAN and edge computing are technically credible.

But successful workload coexistence in a demonstration does not prove that thousands of distributed accelerators will produce an acceptable return.

The business case depends on power, cooling, software licences, local demand, utilisation, redundancy, accelerator-refresh cycles and the revenue available from non-RAN workloads.

No operator should build a national edge-GPU estate before it can identify the anchor workloads that will keep it busy.

Infrastructure should follow repeatable demand, not several years of hope.

Open interfaces can still hide closed intelligence

Open RAN may open the hardware and still leave the operator dependent on one supplier’s telemetry schema, feature store, RIC, model lifecycle, orchestration platform and application marketplace.

That would be a deeper form of lock-in because it controls the path between insight and action.

Every major procurement should therefore separate five assets:

Telemetry. Models. Policy. Execution interfaces. Compute.

A supplier may provide several of them. The operator must retain access, auditability and a realistic substitution path for each.

Physical AI is where the network starts to matter again

Physical AI is often reduced to humanoid robots. That is far too narrow.

It includes any system that senses the physical world, interprets what is happening and triggers an action:

Cameras detecting unsafe behaviour.

Drones inspecting hazardous equipment.

Autonomous vehicles moving through mines and ports.

Robots coordinating across a factory.

Utility systems detecting and isolating dangerous conditions.

Clinical systems recognising deterioration and initiating a response.

The full chain is:

Sense → Infer → Authorise → Act → Verify

From working with large-scale IoT, video and safety platforms, I have learned that the model is usually only one component of the problem.

The failure often happens somewhere else: a sensor is unavailable, video arrives too late, the model is running in the wrong location, the workflow reaches the wrong person, or nobody verifies that the required action actually happened.

This is why mission-critical AI should be the entry point.

Physical AI is a broad market. Mission-critical AI narrows it to processes where a delayed, missed or incorrect action has a measurable cost in safety, downtime, output or compliance.

That creates a willingness to pay for assurance, not merely intelligence.

The edge is not a place. It is a scheduling decision.

The industry often treats physical AI as a contest between device intelligence, telco edge infrastructure and the cloud.

That is the wrong framing.

Qualcomm’s direction is strategically important because it shows how much perception, reasoning and control will remain on the device. Safety functions, immediate machine response and basic autonomy cannot disappear when a wide-area connection degrades.

At the same time, larger multimodal models, multi-camera fusion and fleet coordination may need premises or metro-edge compute. Training, simulation and long-horizon optimisation will often remain centralised.

Qualcomm’s 6G architecture explicitly describes intelligence distributed across the device, RAN, core, edge and cloud, with compute placement driven by latency, privacy, reliability, power and application context. It also places greater emphasis on uplink performance, cell-edge coverage and network adaptation to application intent.

AI-RAN approaches the same continuum from the other direction: putting accelerated compute into network locations so that RAN and external workloads can share infrastructure.

Both directions can be right.

The conclusion is not that all inference will migrate into the telco edge.

It is that inference will fragment across multiple locations and move as conditions change.

The operator’s strategic asset is therefore not the accelerator.

It is the placement and assurance policy:

Which model runs where.

Which data may leave the premises.

Which traffic receives priority.

What happens if the preferred inference location fails.

And how the application falls back without creating an unsafe physical state.

The product should be an assured control loop

This is the clearest way to turn the strategy into something an enterprise can buy.

The service should make one commercial promise:

The right action will be completed within the required time, under defined operating and failure conditions.

That requires four reusable capabilities.

Placement

A workload-placement engine chooses between the device, premises edge, metro edge, telco cloud and approved external cloud.

The decision should consider model size, deadline, jitter, radio conditions, data sensitivity, power, hardware availability and cost.

This cannot be a one-time deployment decision. The placement may need to change while the application is running.

Network intent

The application should communicate what it needs from the network.

Not “give me 100 Mbps.”

More useful intents would be:

  • This video stream supports collision avoidance.
  • This inference must be completed within 80 milliseconds.
  • This workload cannot leave the facility.
  • This control loop must survive loss of the public cloud.
  • This emergency event takes precedence over routine telemetry.

The operator then maps that intent to uplink capacity, QoS, traffic prioritisation, local breakout, network slicing, redundant paths or non-terrestrial backup.

Authority

Every model and agent receives a defined identity and decision scope.

The authority layer specifies which action is allowed, over what assets, at which confidence, with what maximum impact and under which human-approval conditions.

Low-risk actions may be automatic.

Medium-risk actions may use canary execution and automatic rollback.

High-impact actions remain human-approved.

A high model score should never become an unrestricted production credential.

Outcome assurance

The system must observe the whole chain, not merely network latency or GPU availability.

The commercial SLA clock starts when the sensor detects an event.

It stops when the action is completed and verified.

For a factory, that may mean detecting a defect and stopping the correct production stage.

For a utility, it may mean recognising a dangerous condition and isolating the affected asset.

For public safety, it may mean detecting an incident, validating it and initiating the correct response.

For healthcare, it may mean identifying deterioration and successfully escalating it to the appropriate care team.

I would call this a perception-to-action SLA.

The customer is not purchasing low latency in isolation. It is purchasing confidence that a critical operation will complete on time.

What the telco owns and what it does not need to own

The operator does not need to manufacture the robot, build every industry model or write every control application.

It should own:

  • the enterprise service contract;
  • inference-placement orchestration;
  • network and application policy;
  • identity and authority management;
  • end-to-end observability;
  • safe fallback and recovery;
  • and the operational SLA.

Device manufacturers, chipset companies, model providers, industrial software vendors and systems integrators can remain part of the solution.

But they should connect through the operator’s control and assurance layer rather than each creating a separate vertical stack.

The commercial model should reflect this.

Charge a recurring platform fee per site, fleet or critical workflow.

Add an assurance tier based on deadline, availability, locality, recovery and data residency.

Add managed operations where the operator takes 24/7 responsibility.

For selected use cases, add an outcome-linked component based on uptime, avoided downtime or verified operational performance.

That is a better business than reselling GPU hours with a telecom logo.

A CEO’s 24-month agenda

1. Select two control loops, not ten sectors

Do not begin with “manufacturing” or “smart cities.” Those categories are too broad.

Choose two precise workflows where:

  • failure or delay has a high cost;
  • the workload is sensor- or uplink-intensive;
  • local processing or data residency matters;
  • the same architecture can be repeated across sites;
  • and the customer is prepared to pay for assurance.

A machine-vision defect-to-line-stop loop is a product candidate.

“AI for manufacturing” is not.

A drone-inspection-to-maintenance-order loop is a product candidate.

“Drones over 5G” is not.

2. Build the minimum reusable control platform

During the first year, connect one device class, one premises-edge configuration, one metro-edge tier and one cloud environment.

Add the application-to-network API, workload placement, model registry, agent identity, policy enforcement, replay or digital-twin testing, canary execution, rollback and end-to-end outcome telemetry.

Do not attempt to create a perfect digital twin of the whole network or support every accelerator from day one.

Build only what the first two paid use cases require, but build those capabilities so that they can be reused.

3. Put paid lighthouse services into production

Between months 12 and 18, move from a technology pilot to an operational contract.

The contract should state:

  • the event-to-action deadline;
  • the availability target;
  • where data and models may run;
  • the local safe mode if connectivity or compute fails;
  • the recovery-time commitment;
  • and how successful action will be verified.

A free demonstration proves technical interest.

A paying customer accepting an operational SLA proves a business.

4. Scale infrastructure only after proving unit economics

Between months 18 and 24, replicate the control loop across multiple sites.

Then measure accelerator utilisation, power, cloud egress, support cost, deadline compliance, safe fallback, cost per authorised action and recurring gross margin.

Only after those numbers work should the operator expand its distributed compute estate.

At least one early use case should be stopped. A programme in which every pilot survives is not being governed rigorously enough. The attached research lays out the same progression from foundation and bounded production loops to paid vertical deployments and then scale.

The board does not need a dashboard containing fifty AI metrics.

It needs six:

Deadline attainment.

Successful verified actions.

Safe fallback performance.

Avoided incidents or downtime.

Cost per authorised decision.

Recurring gross margin.

This is the discipline behind measuring “value per authorised decision”: the value created after including compute, software, integration, supervision, energy, failures and risk – not before.

The decision

The strategic choice is not between remaining a telco and becoming an AI company.

It is between sitting outside the control loop as a supplier of bandwidth and compute, or sitting inside it as the party responsible for placing intelligence, prioritising traffic, authorising action and proving the result.

Within 24 months, a CEO should be able to answer one question:

Is an enterprise paying us to guarantee a critical physical workflow – not simply to connect it or host its model?

The model may run on a Qualcomm-powered device, an enterprise server, a metro-edge facility or a hyperscaler cloud.

That should not matter to the customer.

The operator should still be able to guarantee where it runs, how it is connected, what it is allowed to do, how it fails safely and whether the required action occurred.

If the answer is yes, the telco has created a new position in the AI value chain.

If the answer is no, the AI strategy is still an internal transformation programme or an infrastructure bet.

The winning telco will not own every model.

It will own the point at which intelligence becomes action.

That is where accountability sits. And that is where the margin can sit too.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top