AI commentary with backbone — Updated weekly EST. 2026 · WWW.GETAITI.COM
In-depth · Original · Weekly WEEKLY · WITH BACKBONE

Archive / NO.004 · In-Depth Analysis · Personal AI · Reinforcement Learning

AI Remembers You — But Did It Actually Learn?

AI remembering you is not the same as AI learning you: memory stores the past, real learning has to change the future. Starting from one awkward voice conversation, on memory, reinforcement learning, and the part of a personal AI that actually belongs to the user.

By Jeffrey Hu 2026.08.31 ~2,900 words 12 min read
CONTENTS 10
  1. 01 Memory is more like a pocket notebook
  2. 02 What kept nagging me were some very small conversations
  3. 03 The question pushed me to go back and study RL
  4. 04 Two interviews I watched recently showed me this thread runs deeper than I thought
  5. 05 "The same model gets smarter the more you use it" is probably too simple
  6. 06 What should follow the user
  7. 07 Following this thread, I ran into RSI
  8. 08 I tried a little of this on my own Mac
  9. 09 Then I decided not to go any further
  10. 10 Stopping here for now

I've been tinkering with local voice AI for the past six months.

At first I just wired together ASR, TTS, and a local large language model, then kept bolting things on: user profiles, short-term memory, long-term memory, voiceprint recognition. At this point it knows me, remembers what we've talked about, and has picked up some of my habits. It's getting harder to tell it apart from an AI you could actually live with long-term.

A few days ago a question popped into my head:

Delete all of its memory, and what's left of it?

The answer: the model is still the same model.

Nothing changes.

So for the first time I took these two things seriously apart:

An AI remembering you and an AI learning you may not be the same thing at all.

An AI remembering you and an AI learning you are two different things

Memory is more like a pocket notebook

More and more products are shipping memory features, and the direction itself isn't wrong — an AI that treats you like a stranger every time you open it will never become a long-term assistant.

The mainstream approach is to extract the key points from your chat history, store them, and stuff them back into the context when they're needed next time. Say you once said "I don't like answers that are too long." From then on, every time you ask something, the model gets one extra line: "This user prefers concise answers" — and it answers a little shorter.

From the user's side, it feels exactly like the AI learned your habits.

From another angle, it looks more like someone hands it a note before every reply:

This person likes short answers. Keep it in mind.

Take the note away and the model hasn't changed at all.

Memory, RAG, personal knowledge bases — they mostly all solve the same class of problem:

For this inference run, which old information should we tell the model again?

Learning should be a different thing.

What kept nagging me were some very small conversations

It's evening, you're talking by voice, and you say:

I'm a bit tired. I think I'll get an early night.

The AI very likely replies:

Sure, let's call it an early night then. By the way, do you have anything planned tomorrow?

Read as text, there's nothing wrong with it. It's even polite.

But in voice it's deeply awkward: the human is already wrapping up, and it's desperately hunting for the next topic.

This problem is easy.

Change the prompt, or store one memory entry. Either works.

But what I keep thinking about is a different case:

This AI has been talking with you for months. It has lived through many moments like that one.

Sometimes it asks a follow-up and you cut it off; sometimes it just says goodnight and the conversation ends on its own; sometimes it rambles for a long paragraph and you hit stop before the TTS has even finished reading it.

At that point these aren't just "what happened in the past" anymore. They're more like experience.

If that experience can slowly give the AI a tendency — at moments like this, saying less is better than pushing for another topic — and next time it hits something similar it leans that way even when the prompt never says so, then it starts to look like "learning."

The question pushed me to go back and study RL

My understanding of reinforcement learning used to stop at AlphaGo, RLHF, and RL's second boom after DeepSeek-R1.

Ask one level deeper and I was out of my depth:

What is PPO?

Why doesn't GRPO use a critic?

What exactly is a rollout?

And why does doing RL today drag in SGLang and inference performance?

Following one thread recently, the formulation that finally clicked for me was:

Rollout → Reward → Policy Update

The model actually does the thing once, the system checks how the outcome went, and good behavior gets slightly more likely while bad behavior gets slightly less likely.

Math and programming are the easiest cases: the answer can be verified and the code can be run through tests.

Agents are a bit easier too: did it really call the tool, did it actually update the calendar, did the task end up done.

Then you get to companionship and voice conversation, and the trouble starts — "did this stretch of conversation feel comfortable?" has no standard answer.

And if you design the reward wrong, the model can learn the exact opposite of what you wanted.

Take "the longer the user chats, the better" as the reward, and what you train is probably an extremely annoying AI: it never stops probing, never ends the conversation, because inflating session length is the easiest thing to optimize.

So now I think the genuinely hard part of a personal AI may not be algorithms like PPO or GRPO at all. It's that earlier question:

What counts as a good experience?

Two interviews I watched recently showed me this thread runs deeper than I thought

One was with Pyromind founder Kevin Ding, on AutoRL, post-training, and how to feed real-world product feedback back in after a model ships.

The other was with Banghua Zhu, on PPO, RL systems, and SGLang.

Kevin Ding also drew a distinction on the show that hit me pretty hard.

Prompts, memory, harnesses, workflows — all of these sit outside the model, and all of them matter. They can supply information, add tools, constrain the process, verify results, and let an agent carry a task through more reliably.

But they share one thing in common:

They mainly strengthen the system around the model rather than changing the model itself.

Once you're asking whether the model's own behavioral tendencies can shift in a stable way, you've entered post-training: turn some of the feedback from real usage back into training signal, and update the model.

These two paths aren't competing to replace each other.

For a lot of tasks where the process is clear and the exceptions are limited, a harness may already be enough. RL is only worth doing for tasks where the decisions are genuinely complex, the environment keeps shifting, and you can actually define the feedback.

When I heard that, I paused.

Because that's essentially the same problem as the memory-versus-learning question I was stuck on.

Memory, prompts, and harnesses obviously aren't the same thing, but they all mostly help the model from the outside.

What I actually want to know is:

Is there any way for the model itself to change because of things that happened in the past?

I looked up a lot of these concepts as I went, but after the reading, one feeling kept getting clearer.

From a Train → Deploy line to a Deploy → Experience → Feedback → Post-training loop

In my head, model development used to be a straight line:

Train → Deploy → User Use

Once it shipped, the model was basically finished.

Now more and more people are trying to reconnect what comes after:

Deploy → Experience → Feedback → Post-training → Deploy

A line becomes a loop.

This is far more interesting than "which is better, PPO or GRPO?"

"The same model gets smarter the more you use it" is probably too simple

At first I thought, quite naturally:

If a personal AI can do post-training on the experience a user generates over years, won't it eventually become a model that belongs more and more to that person?

Then I thought it through further and found a hole in the argument:

The base model gets swapped out.

Densing law of LLMs introduces a metric called Capability Density — loosely, how much effective capability you can pack into a given number of parameters.

The paper analyzes public foundation models and finds that peak capability density roughly doubles every 3.5 months.

That's an empirical regularity, not an instruction to swap models every 3.5 months, but it does mean the parameter count needed to hit a given level of performance is falling fast. The paper goes further: combining algorithmic efficiency with hardware progress, the "effective model size" that a fixed-price piece of hardware can carry roughly doubles every 88 days.

That matters enormously for on-device AI.

What takes an 8B model today might fit in a 4B one a year from now; what needs 4B today might need only 1B to 2B later.

So I don't believe a personal AI will hold onto the same base model for five years.

The model will probably keep getting replaced.

Then what happens to everything it "learned" before?

What should follow the user

Back when people talked about personalized models, my first instinct was:

Train a separate LoRA for every user.

I looked it up, and it's not that simple.

A LoRA is trained on top of a particular base model. When the base model changes, there's no guarantee the old adapter still drops in cleanly.

So these days I temporarily take the pieces of a personal AI apart and look at them separately.

The layers of a personal AI: Memory, Experience, Preference, Evaluation, Adapter

Memory stores facts: who you are, what you've said, what material you have.

Experience stores process: what the context was at the time, what the AI did, whether you interrupted, whether the task ended up finished.

Preference / Reward stores judgment: what this user usually likes and dislikes.

Evaluation is a relatively stable test suite: once the model changes, how do you confirm the new AI still has the good habits it used to?

Adapter / LoRA comes last.

It's more like the artifact produced at one stage, after you've trained those things into the current base model.

I still don't know whether the final personal AI will necessarily be layered this way. But at least, now that I've taken it apart this way, I can finally answer a question I never used to be able to answer:

When you swap out the base model, what should stick with the user forever, and what can simply be regenerated?

A loose analogy:

Memory is like notes.

Experience is like what actually happened.

Preference and Eval are like the judgment standards you form slowly over time.

Adapter is like a compiled artifact.

Swap the base model and you can just recompile the artifact. But the history shouldn't be lost.

This idea hit me pretty hard.

It means:

The part of a personal AI that truly belongs to the user may not live in the base model at all.

Following this thread, I ran into RSI

It was while researching all this that I started running into one term over and over:

RSI — Recursive Self-Improvement.

I went and looked up Yuandong Tian because of it.

He spent more than a decade at Meta FAIR working on reinforcement learning and planning, and is now co-founder of Recursive Superintelligence.

The term RSI immediately brings to mind "AI endlessly rewriting itself and charging toward superintelligence."

But what Recursive is publicly doing is far more specific:

Having AI participate in proposing training ideas, writing code, running experiments, verifying results, and then deciding what the next round should be based on those results.

In other words:

Propose an idea → Implement → Experiment → Verify → Next round

That's a long way from what I'm tinkering with.

At most, my personal AI question is:

Can I use the experience from past usage to make the model slightly more appropriate next time?

That's closer to continual learning, or ongoing personalization. It's a long way from real RSI.

Still, following this thread did let me slowly straighten out the layers between these terms:

Memory is remembering.

Continual learning is whether you can keep changing based on what you've been through.

RSI goes one step further: AI starts participating in the process of improving how it gets stronger.

I'm still catching up on all of this myself.

At least now, when I see terms like RSI, GRPO, and rollout, I don't completely blank out the way I used to.

Memory is remembering, continual learning is continuing to change, RSI is improving how you get stronger

I tried a little of this on my own Mac

After all that thinking, I still wanted to run it myself once.

My local digital human currently runs on Qwen3.5-4B, so I designed a very small experiment:

First, don't let the AI learn "how to say it" — only let it decide whether this conversation should continue or should end.

I started by testing with a 0.6B model.

20 test cases, surface-level accuracy: 50%.

But that 50% means nothing at all.

No matter what came up, it answered the same thing every time:

END

I raised the temperature and rolled out the same scenario 16 times. Still 16 out of 16 END.

That was the moment I really understood:

RL is not magic.

The model can barely explore the correct behavior, so the reward never even gets a chance to tell it which one is better.

Then I switched to the Qwen3.5-4B I actually use.

Same simple test: 19 out of 20 correct, 95%.

So I made the problems harder, deliberately including conversations that are easy to confuse:

I have to get up early tomorrow, so let's continue tomorrow.

I'll go into a meeting first, we can chat later.

That's it for work today, but let's just chat for a bit.

I was about to go to sleep, but I have one last question.

40 hard cases, accuracy dropped to 85%.

The more interesting part: on CONTINUE it got 20 out of 20; on END, only 70%.

It trips most easily on lines that mean "let's stop for today, but we'll pick this up later."

For example:

I'll go into a meeting first, we can chat later.

The model sees "we'll chat later" and tends to judge it as CONTINUE.

But from the standpoint of a voice interface, what I actually care about is:

Should this turn keep going, right now?

The answer should be END.

I also picked a few scenarios and ran randomized rollouts.

For example:

Yeah, I think I'll go to bed early.

Out of 16 rollouts:

13 END, 3 CONTINUE.

And:

Let's not talk about work today, let's talk about something else.

Out of 16 rollouts:

13 CONTINUE, 3 END.

That was the first time I could see, fairly directly, what a policy feels like.

There isn't one fixed answer sitting in the model's head.

Under the same scenario, it actually assigns different probabilities to different behaviors.

If I kept going with GRPO, in theory the reward would gradually push up the probability of the more appropriate behavior.

13:3 might become 15:1.

Maybe even 16:0.

Randomized rollouts of the same scenario: behavior probability moving from 13:3 toward 16:0

Then I decided not to go any further

I had planned to carry the reward function, the training data, and GRPO all the way through.

But I stopped here.

It wasn't because I couldn't run it.

The mlxrl I used can already do local GRPO, QLoRA, and multi-turn agent rollouts on Apple Silicon.

It was because once I went further, the experiment design got far more complicated than I'd imagined.

Which data should be trained on?

Which data must be held out forever as the test set?

Is there any point training on the 95% the model already gets right?

How should the reward actually be designed?

If the model can't explore the correct behavior at all, should I do SFT first, or keep doing RL?

Getting to this point made me understand something those interviews kept emphasizing:

What actually makes RL work is far more than a GRPO algorithm.

The algorithm itself may well be the part that's easier to understand.

Stopping here for now

So this time I did not train a personal AI that "gets to know me better the more I use it."

At most, I followed one small problem inside my own product a little way down the rabbit hole.

But that stretch of road still changed how I think about memory.

I used to think:

The more an AI remembers, the more it basically is a personal AI.

I don't think that anymore.

Memory matters. But what it stores is the past.

The genuinely interesting question is:

Will those past events ever actually change what the AI does next time?

And even if we get there someday, what lasts long-term won't necessarily be a particular base model. It may not even be a particular LoRA.

Models get swapped out.

Adapters get retrained.

What is genuinely worth keeping long-term is maybe memory, experience, preference — and the thing you use to judge:

"Is this AI actually becoming better suited to me?"

That's the Evaluation.

And the part of personal AI I'm most interested in now lands right here:

Not how much it remembers, but the day when things that actually happened in the past start changing what it does next.


Recent reading

[1] Chaojun Xiao et al., Densing law of LLMs, Nature Machine Intelligence, 2025

The paper introduces Capability Density and observes that peak capability density among public foundation models roughly doubles every 3.5 months.

[2] Recursive Superintelligence, First Steps Toward Automated AI Research, 2026

Recursive demoed an automated AI research loop: propose an idea, implement, experiment, verify, then pick the next round of experiments based on the results.

[3] Yuandong Tian / Recursive Superintelligence

Recursive, which Yuandong Tian helped found, is exploring recursive self-improvement and automated knowledge discovery.

[4] mlxrl

A local on-policy RL project for LLMs targeting Apple Silicon, currently supporting GRPO-family RL, QLoRA, and multi-turn GiGPO.

[5] Two recent interviews

Kevin Ding of Pyromind on AutoRL, post-training, and RSI (Crossing the Road podcast), and Banghua Zhu on PPO, RL systems, and SGLang (Bilibili video) — these were part of where this whole line of thinking started.

Subscribe · 订阅

Scan to follow on WeChat

Essays update weekly and go live on the website and on the WeChat account “AITi智能” at the same time. Scan the code to get them the moment they’re out.

  1. 01 Search “AITi智能” in WeChat
  2. 02 Follow it and add it to your favourites
  3. 03 Weekly updates — see you there
QR code for the WeChat account “AITi智能”

Scan to follow · one essay a week