Archive / NO.003 · In-Depth Analysis · On-Device AI · Local AI
From "It Runs" to "It Works": What's Still Missing in Local AI
Running isn't the same as being good to use — models, quantization, inference frameworks, speculative decoding, and agent harnesses are all maturing at once, so local AI's inflection point isn't one magic day but an entire stack crossing the usability line.
CONTENTS 13
- 01 The Other Track Beyond Scaling Law
- 02 Notes From Running Qwen3.8-27B
- 03 There's a Lot More Than the Model
- 04 Is Local Always Faster?
- 05 Now, Cost
- 06 Privacy Isn't Simple Either
- 07 DFlash 2 and Inference Acceleration
- 08 The Agent Harness Layer
- 09 Local and Cloud Aren't Either/Or
- 10 The Inflection Point, as I See It
- 11 What This Means for People Building AI Products
- 12 Last
- 13 References
Running large language models on your own PC stopped being news a long time ago. 7B, 14B, 32B — as long as the hardware can take it, someone has already done it. What made me take another serious look recently wasn't "some 27B model finally fit onto a personal computer." It was a different shift: local models are moving from "can run" to "can do actual work."
I recently ran Qwen3.8-27B on my own Mac, and both its speed and its capability surprised me. But one thing has to be cleared up first: for most personal computers today, 27B is still no light lift. You need enough unified memory or VRAM, enough memory bandwidth, and then you have to wrestle with quantization, inference framework, context length — even whether to turn on inference acceleration like MTP or DFlash 2.
Put plainly, what changed today isn't "your PC suddenly turned into a data center." It's that the model, the hardware, the inference framework, the whole local-AI software stack — several things all moving forward at the same time. That's what I think local AI is genuinely worth watching for.
The Other Track Beyond Scaling Law
For the past few years, large language models have basically followed Scaling Law: more data, bigger models, more compute, in exchange for stronger capability. From GPT, Gemini, and Claude to Qwen and DeepSeek, the ceiling on cloud models keeps getting pushed higher — capabilities that felt remote two or three years ago are now everyday tools. Scaling Law hasn't failed, and cloud models will most likely keep getting bigger and stronger.
But the other side is getting just as obvious: GPU clusters, data centers, power, cooling, networking, plus the token cost of every single call, are all rising along the way. Cloud AI can't absorb every last bit of compute without limit.
So I've been watching a different road: can a smaller model do work that previously required a bigger one?
A while back, Tsinghua University and OpenBMB proposed a concept called Densing Law, the law of density. It doesn't care how large a model can get; it only cares about how much capability is packed into each unit of parameters. The metric is called Capability Density. They analyzed a batch of open-source models and concluded that over the three years from 2023 to 2025, the maximum capability density of large language models roughly doubled every 3.5 months.
If that curve keeps holding, the parameter count needed to reach a given capability keeps shrinking. Work that takes tens of billions of parameters to do well today might need only a dozen billion — or even a few billion — in the future.
That sentence sounds like a prophecy, but I happen to have a ready-made example on hand. When I wrote about the local digital human last issue, that system — speech recognition, conversation, voice, visual presence, memory, real-time interruption — added up to a brain that was only a 4B model, plus two 0.6B models, one for speech recognition and one for speech synthesis. Fewer than 6B parameters in total, and it supported a companion that could talk in real time, be interrupted, hold memory, and have a face. I measured it in the previous issue: the 4B went from the end of your sentence to the start of its reply in a little over 100 milliseconds, at a generation speed of 160 token/s.
Two or three years ago, real-time voice conversation was basically a cloud-exclusive affair, never mind hanging a digital human and a memory system off it. Now it runs entirely inside a MacBook. That's not because one model suddenly got smarter; it's because the capability-density curve is doing its work: the parameter count needed for the same job keeps shrinking.
It resembles Moore's Law, but it isn't pushing the same thing: Moore's Law pushes hardware density, Densing Law talks about model capability density. Stack those two together — compute on endpoints and edge devices keeps climbing, while the models themselves keep getting "small but strong" — and the part of AI that used to be locked inside data centers naturally starts moving next to the user.
Notes From Running Qwen3.8-27B
I'm not using Qwen3.8-27B to prove that "LLMs can finally run locally" — that happened long ago. What's interesting to me is that not long after this model shipped, the community already sprouted a pile of optimized builds for different hardware and different needs.
Xueyu collected a few of them on X, and I'm copying them here:
mlx-community/Qwen3.8-27B-4bit + incoai/Qwen3.8-27B-DFlash2: the MLX route on Apple Silicon; in his own testing, adding DFlash 2 roughly doubled TPS;Qwen3.8-27B-MTPLX-Optimized-Speed: native MTP, model about 20.4GB, balancing speed, quality, and memory;Qwen3.8-27B-MTPLX-Optimized-Quality: mostly 8bit, peak memory about 32.7GB, quality-leaning;- Several builds of
Qwen3.8-27B-GGUF: mainly solving llama.cpp, MTP, multimodality, and cross-platform compatibility.
That "TPS doubled" number comes from the author's own device and can't be treated as a gain every piece of hardware will see. But the example itself is highly representative — what people are discussing has already changed. The question used to be "can this model run at all." Now the questions are: MLX or GGUF on the Mac? 4bit or 8bit? Is memory enough? Quality or speed? Native MTP or an external DFlash 2? How large a context? How do you wire it into OpenCode, an agent, or your own service?
That's a different stage now: the model is just the base layer, while deployment, quantization, inference, acceleration, and the application ecosystem around it carry more and more weight.
Of course that doesn't mean the bar is gone — quite the opposite. For a dense model like 27B, the weights alone take 15 to 20 GB after 4bit quantization. Add the KV cache, the context, compute buffers, the vision module, and the Draft Model that speculative decoding needs, and actual memory use goes even higher. So an ordinary 16GB laptop and an Apple Silicon machine with 64GB or 128GB of unified memory are two completely different experiences. Local AI is getting easier to use, but there's still a distance to "any computer can run a 27B."
There's a Lot More Than the Model
In the past, when people looked at local models, they only asked two questions: how many B? How much memory does it eat?
Only after actually building a product did I discover that parameter count alone is nowhere near enough. The same 27B, with different hardware, a different inference framework, or a different quantization, can feel like a completely different product. Apple Silicon has MLX, Core ML, and Metal; Intel has OpenVINO; NVIDIA has CUDA and TensorRT-LLM. On top of that there are routes like llama.cpp, ONNX Runtime, vLLM, and SGLang. Quantization comes in 4bit, 5bit, and 8bit. Below that sit operator fusion, attention optimization, KV cache management, memory layout, and separate optimizations for CPU, GPU, and NPU.
None of those decide whether the model can start up. They decide whether the user wants to keep using it. How long until the first token? How many tokens per second? How much speed do you lose when the context gets longer? Will memory suddenly blow up? With other software running on the machine, does the AI stay smooth? Those questions are far more grounded than an "XX TOPS" number on a spec sheet.
So building a product on local AI is a systems engineering problem:
Model × inference framework × operators × hardware
Miss on any one layer, and the final experience gets discounted.
Is Local Always Faster?
A lot of people pitch on-device AI by leading with low latency. That's true, but it only says half the story.
Running locally does eliminate the time spent uploading data, transmitting it over the network, queueing on a server, and returning the result, so network latency and jitter are far lower. But that doesn't mean local inference is necessarily faster. The cloud might be an entire GPU cluster, while your local machine is a PC with a few dozen TOPS. Make the model bigger, or fail to tune the framework and hardware together, and the time you saved on the network can easily be eaten back by slower local inference — leaving the total time longer.
So these days I'd rather describe local AI's advantage as "more controllable latency." For real-time voice, device control, local search, camera analysis — steady, predictable latency is often worth more than peak speed.
Now, Cost
Cloud models are essentially billed by the meter: tokens in, tokens out, GPUs burning the whole time, so there's inevitably a running cost that gets more expensive the more you use it.
Local runs on a different ledger. Once you've bought the hardware outright, the model lives on the device, and one more run doesn't add another API fee. Electricity and depreciation are still there, so strictly speaking the marginal cost isn't zero — but for individuals and for a lot of small and medium-sized businesses, it's already low enough to ignore.
This directly rewrites the cost model of a lot of AI products. Knowledge bases, voice assistants, image understanding, local retrieval — services that get called many times a day. Run entirely through the cloud, the long-run token cost is substantial. If high-frequency, simple tasks are all handled locally and only the genuinely complex ones are thrown to the cloud, the cost of the whole system looks completely different. So I've always thought the most concrete value of local AI isn't replacing the cloud — it's offloading the cloud.
Privacy Isn't Simple Either
This is the piece I've been turning over a lot recently.
The moment people bring up privacy, they say sensitive data must not go to the cloud, so you have to use a local model. The logic is right, but reality has another problem: local models still have a ceiling today. For complex reasoning, specialized knowledge, and long context, the strongest cloud models are still clearly stronger.
So product design often gets stuck on a balance: chase the strongest capability and you have to send more data to the cloud; keep all the data local and you have to give up part of the capability.
I recently saw two fairly typical approaches, one in the legal domain and one in the medical domain. Neither is a simple "all local" or "all cloud" setup. Instead, the local model first processes the sensitive information — names, ID numbers, phone numbers, addresses, case numbers — identifying and redacting it on-device, then sends the redacted data to a cloud LLM for the complex analysis. When the cloud returns, the local model fills the necessary real data back in, and finally the result goes to the user for review. The flow is very intuitive:
Raw data → local redaction → cloud inference → local fill-in → user review
This architecture is quite representative. On-device AI doesn't have to go head-to-head with the cloud; it can perfectly well be the layer of local processing in front of the cloud. The real thing to solve isn't "can data go to the cloud," but which data can go and in what form. In high-sensitivity scenarios like medical, legal, financial, and enterprise knowledge bases, this kind of device-cloud collaboration will get more and more common.
DFlash 2 and Inference Acceleration
In the past, on-device performance came down to two things: a faster chip, or a smaller model. But work like DFlash and DFlash 2 has shown me a third path: don't change the hardware and don't change the model — change how inference works.
A traditional LLM generates text serially: produce the first token, then the second, then the third, one after another. Speculative decoding takes a small Draft Model and lets it guess the upcoming tokens, with the main model verifying them afterward. Guess accurately and you can accept several tokens in one pass, which means fewer full inference passes on the main model. DFlash takes a step further, using Block Diffusion to let the Draft Model predict an entire block of candidate tokens in parallel; DFlash 2 then goes on to optimize candidate path selection and within-block prediction. For Qwen3.8-27B there is now a dedicated DFlash 2 Draft Model, and a GGUF version has shipped that works alongside llama.cpp.
The interesting part isn't how many points any Benchmark gained, but that it shows local AI performance improvement will come along at least three lines at once: stronger hardware, more efficient models, and smarter inference methods. And these three end up tugging against each other — DFlash 2 itself eats memory, so on some machines generation gets faster and the memory left over for context actually shrinks. So local inference increasingly looks like an engineering trade-off, not a matter of pushing some single number to the maximum.
The Agent Harness Layer
Beyond inference performance, there's another shift I think matters just as much.
Agents have been hot for the past year, but what I'm watching isn't one particular agent product — it's the Harness Engineering behind it. I now prefer to see the model as the "brain" and the Harness as the set of things outside it that actually let it keep completing tasks: tool calls, the file system, search, memory, context management, permission control, task decomposition, retry on failure, result verification — all of that lives in this layer.
Frameworks like Hermes Agent and Pi are moving in this direction increasingly explicitly: the model can be swapped, but the Tools, Skills, Memory, Context, and Agent Loop outside it stay.
That matters especially for local models. Local models are constrained by hardware and can't always use the strongest model. But as long as you wrap a well-designed Harness around one — it can search when it doesn't know, call a tool when it needs real-time information, break complex tasks into steps, retry when it fails, and double-check the result — a mid-sized model's ability to actually finish the task can be a whole notch higher than a single inference pass.
So from here on, evaluating a local AI just by asking "what model do you use" isn't enough. You also have to look at what's hanging off the outside.
Local and Cloud Aren't Either/Or
By this point my attitude is clear: the opportunity for local AI in this round isn't "on-device models will surpass cloud models any day now." That question is far too simple.
More likely it's a layered architecture. The on-device layer closest to the user does real-time interaction, device perception, simple inference, and private data processing. One level up, edge devices carry local knowledge bases, file understanding, multimodal search, long-term memory, enterprise private data, and can also run larger local models. The cloud keeps doing the most complex inference, very large models, multi-model collaboration, and the work that consumes large-scale compute.
Agents and Harnesses string these layers together. When a task comes in, the system has to judge which layer it should run on: can it be done locally, does it need a bigger model, which data must never leave the device, which data can go to the cloud once redacted, and finally which model, which tool, which agent to call.
Further out, it may not just be "device + cloud" but "device + edge + cloud + multiple models + multiple agents." AI will increasingly look like an entire computing system rather than one model.
The Inflection Point, as I See It
If I had to summarize again now, local AI's inflection point probably wasn't one day when a "small enough and strong enough" model suddenly popped out, but a lot of factors crossing the usability line together.
Capability density is rising, chip compute and memory bandwidth are rising, frameworks like MLX, OpenVINO, llama.cpp, and TensorRT are maturing, 4bit/8bit quantization keeps getting better, speculative decoding like MTP and DFlash 2 keeps pushing generation latency down, and Harnesses let small and mid-sized models complete more complex tasks through tools, search, memory, and verification. Only when all of this stacks up does local AI genuinely go from "technical demo" to "product capability." That's also why, when I look at on-device and edge AI now, I no longer stare only at a single chip or a single model — the whole system decides the experience.
What This Means for People Building AI Products
In the past, building an AI application meant that knowing prompts and calling a few APIs was enough to build quite a lot. But once you really enter the territory of local, edge, and cloud collaboration, the knowledge span opens up at once: you need to know some hardware — what CPU, GPU, NPU, VRAM, and memory bandwidth actually affect; you need to know models — the capability and resource requirements of different ones; and you also need inference frameworks, performance optimization, Agent, Tool, Memory, Context, and Harness, plus where the data should live, when to go local, and when to go to the cloud. Beyond that there is multi-model collaboration, multi-agent scheduling, and an entire AI Infra.
That reminds me of a role that's been getting a lot of mention lately: FDE, Forward Deployed Engineer. These people aren't quite like traditional software engineers. They can't just be nailed to one technical point; they have to get into real scenarios and assemble the model, the engineering, the data, the hardware, and the business into one specific problem, solved. I think device-edge-cloud AI will need this kind of person more and more — they don't have to understand model training best, and they don't have to understand chips best, but they do have to know how to assemble these things into a system that genuinely works.
Last
Over the past few years, what the AI industry cared about most was "how much stronger can the model get." Over the next few years, I think another question matters more: with AI this strong, where should it run?
The answer is probably not the cloud, and not the device either, but rather a judgment based on cost, privacy, latency, capability, and scenario, letting the compute land in the most appropriate place.
Scaling Law will keep going, and cloud models will keep getting bigger and stronger. But at the same time, capability density is rising, local hardware is getting stronger, inference frameworks are maturing, quantization and decoding techniques are changing, and Agent Harness is also raising the success rate of small and mid-sized models on real tasks. So my judgment now is this: what local AI truly deserves attention for isn't "some 27B can finally run on a computer," but that we're moving out of the stage of "will the model even fit" and into the stage of "how do we run it faster, steadier, and cheaper — and make it genuinely useful."
Between "it runs" and "it's good to use," what sits in the middle isn't one stronger chip. It's an entire local AI stack that is maturing fast.
One last aside. In the editorial that launched this magazine I wrote that titanium is valuable because of purification — titanium ore is everywhere, but extracting titanium from the ore takes an extremely long process. Densing Law is actually saying the same thing: refine more capability into fewer parameters. What's scarce in this era has never been a bigger model — it's purifying capability into something that ends up in the user's hands.
References
The concepts, models, and frameworks mentioned in this essay all come from the sources below:
-
Tsinghua University / OpenBMB: Densing Law (the law of density)
- Tsinghua University: "The School of Computer Science's Sun Maosong team proposes the 'law of density,' revealing the inner trend of efficient LLM development"
- https://www.tsinghua.edu.cn/info/1175/122675.htm
- The claims in this essay about "capability density" and "roughly doubling every 3.5 months" come mainly from this source.
-
The Densing Law paper
- Han Xu et al., Densing law of LLMs, Nature Machine Intelligence, 2025
- https://www.nature.com/articles/s42256-025-01137-0
- DOI: https://doi.org/10.1038/s42256-025-01137-0
-
Qwen3.8 / Qwen3.8-27B
- Qwen official GitHub
- https://github.com/QwenLM/Qwen3.8
- The parts about running Qwen3.8-27B locally combine my own hands-on experience.
-
Xueyu: a roundup of community-optimized Qwen3.8-27B builds
- https://x.com/xueyu1125/status/2091493180730687807?s=46
- The MLX 4bit + DFlash 2 combo, the MTP-optimized builds, the roughly 20.4GB / 32.7GB memory footprints, and the author's self-measured TPS gains in this essay all come from that post. Actual speed depends on the device, context, and quantization method, so it can't be treated as a uniform Benchmark.
-
DFlash: Block Diffusion for Flash Speculative Decoding
- Jian Chen, Yesheng Liang, Zhijian Liu, ICML 2026
- https://arxiv.org/abs/2602.06036
- Used to understand how DFlash uses Block Diffusion to draft in parallel and speed up Speculative Decoding.
-
DFlash 2
- Inco AI: DFlash 2: Keep Drafting Parallel
- https://inco.ai/blog/dflash2/
- Qwen3.8-27B DFlash 2 GGUF: https://huggingface.co/incoai/Qwen3.8-27B-DFlash2-GGUF
- On top of DFlash's parallel drafting, DFlash 2 keeps optimizing candidate path selection and within-block prediction, and there is now a dedicated Draft Model for Qwen3.8-27B.
-
Qwen3.8-27B local deployment and quantization examples
- https://huggingface.co/Blackfrost-AI/Qwen3.8-27B-ABLITERATED-GGUF
- That page lists the different quantization sizes from Q2_K to Q8_0, and also notes that actual runtime memory has to account for context state, compute Buffers, the vision module, the Draft Model, and other extra overhead.
-
Hermes Agent
- Nous Research official GitHub
- https://github.com/NousResearch/hermes-agent
- The discussion of Agent Harness, Skills, Memory, MCP, and Agent Loop partly draws on Hermes Agent's public architecture and documentation.
-
Pi Coding Agent
- https://pi.dev/
- Pi calls itself a "minimal agent harness" and supports extending agent capabilities through Extensions, Skills, and Prompt Templates.
-
Video: How far away is the on-device AI inflection point?
- https://www.youtube.com/watch?v=MMy7RJtDxoI&t=66s
- A video discussion running in the same direction as this essay: the industrial inflection point of AI's move from cloud to device, how phones, cars, and personal devices evolve into "personal agents," and which bottlenecks on-device models still have to break through in compute, energy use, inference capability, and commercialization.
Subscribe · 订阅
Scan to follow on WeChat
Essays update weekly and go live on the website and on the WeChat account “AITi智能” at the same time. Scan the code to get them the moment they’re out.
- 01 Search “AITi智能” in WeChat
- 02 Follow it and add it to your favourites
- 03 Weekly updates — see you there
Scan to follow · one essay a week