Why optimize the same GPU instead of simply buying more compute?
LG Uplus is jointly researching token optimization for server GPU environments with AI model optimization company OptAI.
LG Uplus says the ongoing early-stage research improved the model's computation for real service environments and increased the number of tokens the same GPU could process by up to four times.
That does not mean company-wide AI costs have fallen by 75%, nor that every model will see a fourfold gain. It is an early result. What matters is the direction: AI competition is expanding from bigger models toward extracting more useful work from the same compute.
User growth turns model performance into an operating-cost problem
At the beginning of an AI service, answer quality and model performance dominate the discussion.
Once usage grows, however, repeated inference becomes GPU consumption, electricity and infrastructure cost. LG Uplus explicitly frames efficiency as increasingly important because expanding AI services raise GPU and power requirements.
The operating question therefore grows from 'How good is the model?' to 'Can this model be called millions of times economically?'
Throughput per GPU can become a business metric
Tokens are the basic units an AI model processes when understanding prompts and generating responses. More token throughput on the same GPU creates room to serve more requests without increasing compute at the same rate.
In manufacturing terms, it resembles raising output per machine without adding another production line.
For AI services, tokens per GPU, cost per request, latency and utilization increasingly connect technical architecture to business economics.
LG Uplus had already tested the same logic on smartphones
The server project follows earlier work by LG Uplus, LG AI Research and OptAI on an EXAONE 3.5-based small language model for smartphone NPUs.
According to LG Uplus, the NPU-based model maintained a similar performance level to the prior CPU approach while reducing power consumption by 78% and model size by 82%.
The earlier constraint was a small device with limited power and memory. The new constraint is large-scale server GPU operation.
Optimization is expanding from device constraints to server constraints
On-device AI must fit within memory, battery and limited local compute. Server AI must handle GPU availability, power, inference cost and large volumes of concurrent requests.
The environments differ, but the optimization question is similar: how much resource can be removed while preserving the service quality that matters?
The sequence can be read as Device Constraint → Model Compression → Server Constraint → Token Optimization.
Build Better Models and Run Models Better are different capabilities
Model competition makes benchmarks, accuracy and reasoning performance highly visible. Production AI introduces another stack: compression, quantization, inference optimization, GPU utilization, serving architecture and latency.
If two services deliver comparable quality but one consumes materially more GPU resources per request, that gap compounds as user volume rises.
The capability to build a stronger model and the capability to operate a model more efficiently can therefore become separate sources of advantage.
The AI talent market expands downstream into systems efficiency
Researchers and model engineers remain important, but scaled AI services also need people who improve inference, compression, quantization, GPU utilization, serving infrastructure and latency.
These roles may not create a new foundation model. They make existing models run better under real traffic and real cost constraints.
As AI moves from research projects into persistent services, Run Models Better can become as strategically relevant as Build Better Models.
BANSEOG VIEW | As AI becomes ubiquitous, using less can become an advantage
The up-to-fourfold token-throughput figure is an early result from ongoing server GPU optimization research and should not be generalized to every model or service.
But combined with the earlier on-device efficiency work, the direction is clear: scaling AI requires not only adding GPUs, but increasing the productivity of compute already owned.
Banseog reads the shift as AI Capability → User Scale → Compute Cost → Optimization → Unit Economics. At scale, delivering comparable AI quality with fewer resources can become a competitive capability in its own right.
BANSEOG VIEW
Banseog View — AI Capability → User Scale → Compute Cost → Optimization → Unit Economics
As AI services scale, compute cost per request and GPU productivity matter alongside model quality.
LG Uplus is extending its optimization work from device NPUs to server GPUs.
The talent stack can also expand from model research toward inference optimization, serving and production efficiency.
SOURCES
Primary sources and references
- LG — LG Uplus advances token optimization with OptAI
October 2, 2026. Confirms joint server GPU token-optimization research, an early result of up to 4× more token processing on the same GPU, and plans for phased application to AI services and infrastructure.
- LG — LG Uplus develops EXAONE 3.5-based on-device sLM
September 25, 2025. Confirms NPU-based on-device optimization with 78% lower power consumption and 82% smaller model size versus the previous CPU approach at a similar performance level.
The 'up to 4×' figure is an early result from ongoing server GPU token-optimization research by LG Uplus and OptAI. It does not imply a fourfold improvement or equivalent cost reduction across every model and service. AI Capability → User Scale → Compute Cost → Optimization → Unit Economics is Banseog's analytical frame connecting the two official releases.