That last point is the one I wish more product pitches would admit out loud. Hiding who paid does not hide what you typed. Masking an IP address does not scrub a writing style. A clever local rewrite can still leak a city, a clinic habit, or a travel pattern if the prompt is too helpful. Buterin’s test is interesting less because it claims a finished product and more because it treats privacy as three separate leaks that have to be plugged at the same time.
Why Personal Health Advice Breaks Ordinary AI Privacy
Health coaching sounds innocent until you list what a good coach actually uses. Sleep. Training load. A sore knee. A week of airports. Food you actually eat, not the food you claim to eat. Those details are useful precisely because they are specific. Specificity is also what makes them identifying. A frontier model does not need your legal name if the prompt already contains a rare combination of schedule, symptoms, and cities.
I’ve found that people underestimate this. They think privacy fails when a company stores a file with a name on it. Often it fails earlier, inside the question. “What should I eat after a 14-hour flight into a humid city if my resting heart rate is up and I cannot do impact work for six weeks?” That sentence is already a small dossier. Add a second message the next day and the dossier gets thicker.
Buterin’s stated goal was narrow and, to my eye, honest. Use personal health and travel data for recommendations. Use stronger remote models when local reasoning is not enough. Avoid leaking private information to those remote models. The local system was Alibaba’s Qwen3.8-Flash-Next, described in his notes as Qwen 3.8 Flash. It decided what a remote model needed, then rewrote the request. Personal records stayed with the local side. The remote side received a thinner slice.
You need all three. Content, payment, and the network path. Drop one and the other two start leaking around the edges.
Vitalik Buterin, describing the experiment
He also said the local model ran around 20 to 30 tokens per second, and that comfort would start somewhere above 100. Tor latency landed roughly 10 to 100 times higher than he wanted. Those two numbers matter more than the architecture diagram. A privacy design people abandon because it feels slow is not a privacy design. It is a demo.
The Self-Experiment, Without the Gloss
Strip away the crypto vocabulary and the test is familiar. A person wants advice that knows their body and their calendar. The best general models live somewhere else. Sending the whole notebook feels wrong. Sending nothing makes the advice generic. So an intermediary has to choose.
Buterin said the local model produced diet and exercise recommendations, and that answers coming back from frontier models improved the result. He did not publish the health records, the detailed plan, or an outside check on accuracy. That absence is worth sitting with. A privacy win that cannot be audited for quality is only half a story. Maybe the advice was sharp. Maybe it was plausible and slightly off, which is a known failure mode when context gets stripped. We do not get to see which.
Perhaps the most interesting aspect is the skill file he mentioned. A set of rules telling the local model when to call out, and how to build a request with less identifying material. That is not a cryptographic trick. It is editorial judgment, automated. Get the rules wrong and you either overshare or starve the remote model of the one fact it needed.
What “Private” Actually Has to Cover
A lot of privacy talk collapses into a single switch. On or off. This experiment refuses that. Buterin framed three layers.
- Request content, so the wording and the personal context do not travel intact.
- Payment identity, so the bill does not point back at the person who asked.
- Network traffic, so the ordinary IP address is not sitting next to the prompt.
Hide only the payment and the provider still reads the prompt. Hide only the IP and the writing style can still cluster sessions. Hide only the content and the account that paid can stitch the calls together. The line he used was simple. You need all three.
I would add a fourth leak that no protocol fixes by itself: memory on the user’s side. If the same laptop, the same hours, and the same topic keep showing up, correlation does not require a name. Privacy here is a reduction of linkage, not a magic cloak.
Layer One: Let the Local Model Rewrite the Ask
The first layer is the least glamorous and, in my experience, the one that decides whether the rest matters. The local Qwen model does not merely summarize. It constructs the outbound query. Original wording stays home. Full personal context stays home. The remote model sees a request the local system judged safe enough.
Why rewrite style as well as facts? Because prose is a fingerprint. Short sentences, a habit of listing constraints, a particular way of describing fatigue. Over enough messages, style can tie sessions together even when names are gone. A rewrite that flattens voice is doing quiet work.
There is a cost, and Buterin named it directly. The more careful you are with what leaves the machine, the less help the remote model can give. That sentence should be taped above every privacy product roadmap. Care and usefulness pull against each other. Pretending they do not is how you get either a leaky assistant or a useless one.
Imagine the local model holding a travel week and a training block. A careless handoff might say the user lands in a named city after a named conference, with a named injury from a named clinic. A careful handoff might say an adult traveler has two long-haul segments, limited walking tolerance, and needs meals that travel well for five days. The second prompt is worse for the chef and better for the person. Which cut is right depends on the question. A skill file has to make that call again and again.
What the Local File Is Allowed to Know
Nothing in this design claims the local model is ignorant. It is supposed to know the health notes and the travel notes. The privacy boundary is not “no model sees the data.” It is “the model that sees the data is the one you run.” That distinction gets lost in marketing. Local inference still processes sensitive text. If the laptop is stolen, if logs are careless, if a plugin phones home, the boundary moves.
Qwen3.8-Flash-Next landed as an open-weight foundation model on August 26, with local inference paths through common serving stacks such as vLLM and SGLang. Open weight is not the same as private by default. It means you can run it without shipping every token to the lab that trained it. Weights on disk, prompts in memory, caches on a drive. The operational hygiene still belongs to the person running the box.
Buterin had already been trying local Qwen models before this health test. The newer step is the intermediary role. The local system is not asked to answer every hard question alone. It is asked to be a gate. That is a different job from “run a chatbot offline,” and it is closer to how careful analysts already brief an outside expert. You do not hand over the folder. You hand over the question the folder implies.
Speed Is Part of the Privacy Design
Twenty to thirty tokens a second is usable if you are patient and the answers are short. It is a poor feel for a back-and-forth coach that rewrites, calls out, waits on Tor, then explains the result. Buterin’s threshold, above 100 tokens a second, is a human-factors number more than a benchmark brag. Below that, people start pasting raw notes into the fast remote box “just this once.”
I have watched that failure in ordinary office tools. The secure path is three clicks slower, so the sensitive paragraph goes through the convenient path. Any private AI setup that cannot keep up with a tired evening question will lose to the tab that is already open. Hardware, quantization, and batching are not side quests here. They are the lock on the door.
| Piece of the setup | What it is meant to hide | What it still exposes |
| Local rewrite | Original wording and unused personal context | Whatever the skill file chooses to include |
| zkAPI payment | Link between a funded balance and a single request | The prompt itself, once it reaches the provider |
| Tor path | Ordinary IP address | Timing, and sessions that reuse the same circuit |
| Local speed | Nothing cryptographic | User patience, which decides whether the stack gets used |
Read that last row twice. The table is not only about math. It is about whether a person at 11 p.m. will bother.
Layer Two: Pay Without Pointing at the Ask
The second layer is zkAPI, introduced on October 1 by the Ethereum Foundation, built with the Open Anonymity Project, and running on Ethereum mainnet. The pitch is metered API access without tying each call to an identity. A user funds a private balance. Later they prove that enough value is available, without showing which deposit is paying for this particular request.
Split the roles and the idea gets clearer. The service that checks payment does not need the prompt. The model provider receives the prompt without learning which funded balance is attached. That separation is the whole point of the zero-knowledge payment proof in this context. It is not trying to encrypt the question. It is trying to stop the invoice from becoming a name tag.
Official descriptions of zkAPI are careful on this, and they should be. The upstream provider still sees prompts. Network and timing data can remain visible outside the proof. The October 1 explanation drew the same line. Payment privacy breaks the link between a user and a request. Content privacy and network anonymity need other tools. Reused personal details, writing habits, conversation history, or documents can still connect sessions.
So if someone tells you zkAPI makes AI private, they have skipped a chapter. It makes payment linkage harder. That is valuable. It is not the whole problem Buterin was poking at with health notes.
How a Private Balance Changes the Habit
Most AI billing today is an account. Log in, get a key, every call wears that key. Even if the vendor promises not to train on your text, the account graph exists. Support tickets, invoices, abuse flags, and subpoenas all know where to look. A prepaid private balance does not erase legal process. It does change the default join between “this person paid” and “this prompt arrived at 6:12.”
There is a consumer version of this fight already, outside crypto. People buy gift cards, use separate emails, or route work questions away from personal accounts. zkAPI is a more formal attempt at the same instinct, with a proof instead of a sticky note. Whether the proof holds up under real traffic, real providers, and real compliance teams is the open question. Mainnet launch is a start, not a finish.
I keep a small skepticism about metered privacy tools that still require a sophisticated client. The people who most need health privacy are not always the people who will run a daemon, fund a balance, and read a threat model. If this stays a power-user stack, it still matters. It just will not be the default path for a clinic portal or a fitness app.
Layer Three: Tor, and Why It Feels Wrong for Chat
Tor is the third layer. It hides the normal IP address from the services receiving the requests. Buterin was blunt about its fit. Tor was not built for the request-by-request unlinking he wants, where separate AI calls would be hard to associate with one another. In testing, latency ran about 10 to 100 times higher than he considered desirable.
That range is wide on purpose. Tor performance moves with circuits, congestion, and what you are asking the far end to do. A short classification might be annoying. A long reasoning call, already slow, becomes a coffee break. For a health coach you might query twice in an evening, maybe you wait. For anything conversational, the wait trains you to stop.
He linked a proposed change in the Ethereum zkAPI repository, pull request 16, still open as of October 4 and not merged. One commit touches seven files. The patch spins up a fresh temporary Tor client when the zkAPI daemon starts, with a new data directory and a new Tor connection. Another command can restart the service so a new network identity appears before a single request or a new conversation.
Read the fine print on identity rotation. A separate client script in the proposal says a fresh server is created for one request or the start of a new conversation. Messages that continue inside the same conversation keep the existing server. They do not automatically get a new Tor identity every turn. That is a reasonable engineering choice and a real privacy limit. A thread is still a thread.
Timeouts Tell You the Network Is the Bottleneck
The patch also lengthens network timeouts, which is the code’s way of admitting Tor is slower. A model-list timeout moves from one minute to three. Other request limits rise from 15 seconds to 60, and from five seconds to 30. If you have ever shipped a client that assumes a snappy API, you know why those numbers change. The old limits would simply fail closed, or fail open in ugly ways, once the onion path entered the picture.
Buterin called Tor one of the weakest parts of the current experiment. I think that is fair, and not a dunk on Tor. Tor is excellent at a job it was shaped for: browsing and a class of hidden services, with circuits that are sticky enough to be usable. Per-request unlinkability for chatty model calls is a different job. You can approximate it by burning circuits. You pay in delay, and you still leak timing if someone is watching both ends carefully.
Rough shape of one careful call: local model reads private notes skill file cuts the prompt zkAPI proves a balance can pay Tor carries the thinner request remote model answers local model folds the answer back into the plan
Every arrow in that list is a place to wait, log, or correlate. The design is only as private as the sloppiest arrow.
What the Remote Model Still Sees
This is the section product pages skip. The privacy stack does not stop a remote provider from reading whatever was deliberately placed in the prompt. Documentation around zkAPI says the upstream side still sees prompts. The Foundation’s launch note said the same thing in plainer language. Payment hiding is not content hiding.
So the local rewrite is not a bonus feature. It is the content control. If the skill file includes a rare medication schedule, a niche sport, and three airport codes, the remote model has that bundle whether or not the invoice is anonymous. Providers can log it. Employees can see it if the policy allows. A breach can spill it. Training opt-outs, where they exist, are a policy choice sitting next to that text, not a mathematical deletion of it.
There is also session glue that no single prompt reveals. Come back tomorrow with a follow-up that only makes sense if you saw yesterday’s answer, and you have built a chain. The Tor patch’s choice to keep one circuit for a conversation matches how people actually chat. It also matches how linkage actually happens.
A Diet Question, Told Two Ways
Consider a concrete rewrite, invented here only to show the cut. The private note might say a 31-year-old in Lisbon, flying to Singapore on Thursday, managing a stress fracture, lactose intolerant, aiming to keep lifting twice a week. The careless prompt ships all of that. The careful prompt might ask for high-protein travel meals without dairy, suitable for someone who cannot run, with long-haul jet lag, no city names, no age if it is not needed, no exact diagnosis if “bone stress, no impact” is enough.
The careful version can still be identifying if it is weird enough. “Bone stress, two lifts, dairy-free, two long flights, humid destination” is a smaller needle, not no needle. Privacy work is often about shrinking the haystack match, not deleting the haystack. Anyone selling certainty here is selling a mood.
Buterin said the request-writing rules still need work for exactly this reason. Remove too much and the frontier model shrugs. Leave too much and the experiment fails its own goal. That tension will not be solved by a single pull request.
Why This Sits Next to Earlier Privacy Warnings
The health test did not appear from nowhere. In April 2025, Buterin argued that rising model capability plus centralized data collection made stronger privacy tools more urgent. You can disagree with the policy taste and still grant the mechanical point. Models are getting better at inferring the unsaid. Companies are getting better at joining files. The combination is a poor match for medical-ish notes typed into a chat box.
Ethereum’s own roadmap talk in August included stronger protocol privacy alongside quantum-resistance work and native rollups. zkAPI is a narrower object than “private Ethereum,” but it rhymes with that direction. Pay for an outside service without painting your address on every call. It is infrastructure for a world where the chain is not the product. The product is an API, and the chain is the receipt that does not snitch.
I do not think every reader needs to care about rollups to care about this experiment. The transferable lesson is the split. Keep the sensitive file local. Pay in a way that does not identify the call. Move the packet so the network address is not your home address. Then accept that the answer quality drops as the prompt gets cleaner. That trade is the adult version of private AI.
What a Normal Person Can Steal from the Design
You do not need a custom daemon to borrow the habit. Most people will never run Qwen next to Tor. The pattern still travels.
- Write the private facts in a note the remote chat never sees.
- Ask a local or even a manual rewrite to produce the smallest useful question.
- Send that question, not the diary.
- Paste the answer back into the private note and decide what to trust.
- Avoid follow-ups that only make sense if the vendor kept yesterday’s secrets.
Step five is the one people skip. A clean first message followed by “use the injury I mentioned” rebuilds the file on their side. If you want the remote model forgetful, you have to stay forgetful in how you refer back.
Separate accounts help a little. Separate networks help a little. Neither replaces the rewrite. I would rather send a dull prompt from my real IP than a vivid medical novel through a VPN and call it private. Dull is underrated.
Where the Advice Can Still Go Wrong
Privacy is not accuracy. A model can respect your redaction and still invent a calorie target, miss a drug interaction you failed to mention, or push a training load that ignores a detail you cut for safety. Buterin did not release an independent evaluation. Anyone treating the output as a clinician is doing something the experiment never claimed.
There is a nastier failure mode. The local model can over-trust the remote answer and write it back into the health note as if it were a measurement. Next week the note looks like data. It is an echo. Good intermediaries should mark outside suggestions as suggestions. I have not seen that discipline described in the public notes, and I would want it before I let a stack edit my own training log.
Health data also has other people in it. A travel companion, a shared household meal plan, a child’s schedule accidentally sitting in the same calendar export. Redaction that thinks only about the account holder will miss the bystanders. That is not a crypto issue. It is a notebook issue.
Performance Numbers, Held Lightly
The figures worth remembering are few, and they should stay labeled as one person’s test.
- Local Qwen3.8-Flash-Next around 20–30 tokens per second on his setup.
- A comfort target above 100 tokens per second.
- Tor latency about 10–100 times higher than desired.
- Timeouts in the proposed client stretched to minutes and to 30–60 seconds.
- Qwen3.8-Flash-Next public release dated August 26.
- zkAPI public introduction dated October 1, on Ethereum mainnet.
- Tor client pull request still open on October 4, not merged.
None of those numbers is a law of physics. A faster machine moves the token rate. A quieter Tor circuit moves the latency. A merged patch moves the defaults. What should stick is the shape. Local inference is the gate. Remote models are the expensive brain. Payment proofs and onion routing are the attempt to stop the brain from learning who knocked.
The Open Patch, and What It Does Not Promise
An open pull request is a proposal, not a feature you can assume in production. PR 16 adds Tor-routed client support and adjusts timeouts so slower paths do not die immediately. It creates a temporary Tor client at daemon start. It offers a restart path for a fresh network identity before a new request or conversation. Continued messages reuse the server.
If you are evaluating this as a user, the unmerged state matters. Behavior can change in review. Defaults can tighten. A “fresh identity per conversation” choice might later become optional, or mandatory, or dropped because operators hated the delay. Treating a screenshot of a branch as a privacy guarantee is how people get surprised.
The same caution applies to zkAPI itself. Mainnet means real value can move. It does not mean every provider you might want has integrated, or that their logging policy matches the proof’s story. The proof can be perfect and the vendor can still store prompts in plaintext. Those are different rooms in the building.
A Useful Way to Judge Similar Demos
Next time a project claims private AI, I would ask a short list before applauding.
- Does the remote model see a rewritten prompt, or the original note?
- Can payment be joined back to the account that asked?
- Is the network path separate from the everyday IP, and for how long does that path stay sticky?
- What happens to logs, caches, and conversation memory on both sides?
- How slow is the private path compared with the leaky one?
- Who checked whether the advice got worse after redaction?
Buterin’s write-up answers several of those and leaves others open. That is better than a landing page with a shield icon. The gaps are part of the information. No published health file. No outside accuracy review. Tor still clumsy for per-call unlinkability. Local speed still short of the comfort line. Rules for what to omit still unfinished.
The more careful you are with information sent remotely, the less assistance the remote model can provide.
That limit is the sentence I would keep. It is not a bug report. It is the price list.
Frontier Models as Consultants, Not Roommates
The metaphor that fits this stack is a consultant behind a screen, not a roommate with a key. You brief them. They do not live in the house. The local model is the colleague who knows the messy context and writes the brief. zkAPI is the accounts department that pays the invoice without stapling your passport to it. Tor is the alley you walk so the office camera does not catch your commute. None of them stop the consultant from remembering the brief.
Once you accept that, the product question changes. Not “how do we make the remote model forget,” which you mostly cannot verify. Rather “how little can we put in the brief and still get a useful memo?” Health and travel are a harsh test because the useful memo wants the messy context. Code review might need less biography. Tax questions might need numbers without narratives. The same three layers will not weigh the same in every domain.
Diet and exercise were a revealing choice exactly because they tempt oversharing. People want the plan to feel personal. Personal is the leak. A generic meal template needs no Tor circuit. The moment you ask for personal, you inherit the whole problem.
What I Would Watch Over the Next Few Months
First, whether the Tor client work merges, and whether operators actually rotate identity between conversations or leave one circuit up for convenience. Second, whether any model provider accepts zkAPI payment without demanding a regular account on the side. A proof that still sits behind a login has not finished the job. Third, whether local speeds on this class of model cross the line where rewriting feels instant. Fourth, whether anyone publishes a redaction eval: same questions, full context versus cut context, scored by a human who knows the hidden file.
That fourth item is the one I care about most. Without it, we can praise the plumbing and still have no idea if the coach got dumber in a dangerous way. Privacy that silently degrades medical-adjacent advice is not a free good. It needs a measurement.
I also want to see boring failure stories. Circuits that died. Proofs that were rejected. Skill files that shipped a city name anyway. Demos that only show the happy path teach less than a week of broken evenings. Buterin’s latency complaint is already more useful than a green checkmark.
A Practical Threat Sketch, Not a Panic
Who actually benefits if this stack works? Someone whose travel and health notes would be awkward in a vendor log. A public figure. A person in a household that shares devices. A patient who does not want a fitness app and a model lab holding the same story. The adversary is not always a spy novel. Sometimes it is a future breach, a subpoena, a bored insider, or a training corpus that was not supposed to include that export.
Who does the stack not protect? Someone who pastes the raw file “to save time.” Someone whose local machine is already compromised. Someone who asks follow-ups that rebuild the file. Someone who needs emergency medical care and should be talking to a clinician, not a rewritten prompt. Tools have a lane. This one is a lane for discretionary advice, not a hospital.
There is a regulatory shadow too, even if the experiment is personal. Health-adjacent text triggers duties in many countries once a company holds it. A user running models at home sits in a different bucket from a startup storing the same notes. zkAPI does not convert a company into a non-holder of prompts. If they receive the text, they hold the text, proof or no proof.
How the Three Layers Fail Alone
Walk through the solo failures, because the slogan “you need all three” is easy to nod at and easy to forget.
Content only. You rewrite beautifully, then pay with the same account you use for everything else, from the same home IP. The provider may not know your diagnosis from the prompt, but they know the customer, the hour, and the fact that this customer suddenly asks careful health-shaped questions. Linkage does not require the secret itself.
Payment only. The invoice is a ghost. The prompt is a memoir. Tor is off. The vendor stores a rich health narrative next to a datacenter IP that still roughly places you, plus timing that matches your other traffic if they can see it. The ghost invoice did not save the memoir.
Network only. You are on Tor, proud of it, and you pasted the clinic letter. The exit node does not know your home address. The model vendor knows the letter. Privacy theater with extra latency.
All three, sloppy skill file. You did the hard infrastructure and the rewrite included the one rare detail that makes you searchable. Infrastructure cannot outvote a bad brief.
That is why the local model’s judgment is the load-bearing wall. The chain proof and the onion path are real, and they are not the wall.
Local Models Are a Habit, Not a Brand
Qwen is the model in this test. It does not have to be the model in yours. The durable idea is an open-weight system fast enough to sit in the loop, instructed to minimize identity in outbound text. Brand loyalty is beside the point. If another local model follows the skill file more reliably, use that one. If a smaller model rewrites faster and a larger one only answers the hard remainder, that split may be saner than asking one model to do both jobs.
Buterin’s earlier local experiments make more sense in this light. You learn the failure modes on low-stakes prompts, then you let the model touch health notes. Skipping the rehearsal is how a skill file ships a surname on day one. I would want a red-team hour where a second person tries to re-identify the outbound prompts before any real file is connected.
Speed again. At 20–30 tokens a second, a careful rewrite of a long brief is a pause you feel. At the hoped-for rate above 100, the pause shrinks toward ordinary typing. The privacy strategy gets more realistic as the pause shrinks, because humans are impatient in a very predictable way.
What This Does Not Say About Ethereum
It is tempting to inflate a personal stack into a chain narrative. Ethereum hosts the payment proof. It does not host your diet. The model still runs wherever the provider runs it. Tor is not an Ethereum feature. The local Qwen process is not a smart contract. Readers who only track price will miss the actual claim, which is modest. A public network can settle a private balance check so an API bill does not have to be a profile.
That modest claim is still worth having. So much “crypto plus AI” chatter is a token stapled to a wrapper. Here the chain piece has a job: unlink payment. The AI piece has a job: answer a thin question. The local piece has a job: decide the thinness. If any of those jobs is fake, the demo collapses into a thread. As described, the jobs are distinct. That is rarer than it should be.
A Closing Look at the Trade
Buterin tried to get personalized diet and exercise guidance from stronger models without handing them the life that made the guidance personal. The method was a local rewrite, a zero-knowledge payment split, and a Tor path, with eyes open about delay and about how much help you lose when you withhold context. The recommendations existed. The underlying records did not go public. The accuracy was not independently scored. The Tor support sat in an open pull request. The local model was fast enough to test and, by his own standard, not fast enough to feel comfortable.
I come away thinking the experiment is useful as a template and incomplete as a product. Useful, because it refuses the single-switch fantasy. Incomplete, because request-by-request unlinkability is still awkward, because speed still nudges people toward the leaky path, and because we cannot see whether the advice survived the redaction. That is a respectable place for a self-experiment to stop. It is a bad place for a startup to start selling shields.
If you take one habit from it, make it the rewrite. Keep the file at home. Send the smallest question that can still be answered. Pay and route in whatever way you actually understand. Then read the answer as a suggestion from a consultant who never saw the house, because if the stack worked, they did not.
And if the private path is so slow that you cheat, count that as a failed test, not as a personal lapse. Privacy that only works for people with spare time is a research note. The next version has to be quick enough that cheating feels pointless. Until then, the honest description is the one this trial already hints at. Three layers, partial protection, real latency, and a standing bargain between secrecy and usefulness that nobody gets to skip.
]]>