Can AI Firms Police Themselves On Safety Rules
A Facebook whistleblower just asked the question markets keep dodging: if AI firms sign a safety pledge, will they honor the spirit or slip around the fence? The answer may decide who gets hurt next.
Financial market analysis from 02/10/2026. Market conditions may have changed since publication.
I kept replaying one line after the interview ended. Not the pledge. Not the photo of chief executives around a table. The line about walking around the end of the fence. If you have ever watched a company treat a rule like a puzzle instead of a promise, you already know why that image sticks. Flexible systems do not need to break a sentence to miss the point. They only need a gap wide enough to step through.
Frances Haugen, the former product manager who brought internal research on teen harm into public view, used that fence metaphor on Friday while talking about the fresh White House agreement on artificial intelligence. Major firms, including the company she once worked for, had just put their names on a document that asks the industry to police itself. Her question was simple, and a little uncomfortable. Will they follow the spirit, or only the words that can be checked by a lawyer?
That is not a side debate for policy wonks. It is a market question, a product question, and, if you use these tools with your family, a personal one. I have found that the gap between a signed page and daily practice is where most tech risk actually lives.
Why A Signed Pledge Is Not The Same As A Safety System
Self-regulation sounds tidy. Companies know their models. They move faster than legislatures. They can update a policy on a Tuesday and ship a patch on a Wednesday. On paper, that speed is an advantage. In practice, speed without an outside witness often becomes speed toward the metric that pays.
Haugen’s point was narrower than a blanket attack on the industry. She said AI companies have been more proactive about safety conversations than social platforms were in earlier years. That is a real shift, and it deserves credit. Engagement with risk, public letters, red-team talk, and calls for outside auditors are not nothing. They are also not a finished system.
The history she described is familiar to anyone who has sat in a compliance review. When the product is flexible, teams can satisfy a narrow written test and still miss the harm the rule was written to prevent. A filter blocks one phrase. A ranking tweak hides one complaint. A dashboard turns green. The underlying incentive stays put.
Often they will follow the exact written statement when we are dealing with very flexible systems.
Frances Haugen, on how tech firms meet formal rules
Flexible is the key word. A social feed is flexible. A chatbot is flexible. An agent that books, writes, recommends, and remembers is even more flexible. You cannot bolt a single sentence onto that kind of machine and call the job done. The machine will find a path you did not name.
The Fence And The Gap Beside It
Picture a pasture fence built to keep cattle off a road. The posts are straight. The wire is tight. Then someone notices the fence stops ten feet before the tree line. Nothing in the contract said the wire had to reach the trees. The cattle find the opening by Thursday.
That is the behavior Haugen flagged. Firms can honor what is expressly prohibited and still go around the end of the fence. The agreement might ban a listed misuse. It might not describe the adjacent misuse that appears once the model is in a new product, a new country, or a new plugin. Narrow compliance is not the same as intent.
Perhaps the most interesting aspect is how ordinary this pattern is. It is not a cartoon villain move. It is what happens when a quarterly target, a launch date, and a vaguely worded pledge share the same calendar. Good people cut corners when the corner is rewarded and the witness is absent. Haugen said as much when she talked about independent auditors.
What The Agreement Actually Asks For
The document signed this week is a voluntary framework tied to the current administration. It asks participating companies to step up on safety, security, and responsible deployment. Public descriptions emphasize cooperation rather than a new statute with fines attached. That design is the whole tension. A pledge can move culture. A pledge without a check can also become a press release.
Haugen’s ask was direct. Companies need to step up and comply with the spirit of what they signed, not only the clauses a reviewer can tick. Spirit is harder to audit than a checkbox. Spirit shows up in what gets shipped when nobody is filming.
- Letter compliance means the banned example is blocked in a demo.
- Spirit compliance means the same harm is hunted in products the demo never showed.
- Letter compliance staffs a policy team for the announcement week.
- Spirit compliance keeps that team funded after the headlines move on.
- Letter compliance treats an audit as a threat.
- Spirit compliance treats an audit as the only way angels stay honest.
I do not think every signer is looking for a loophole. Some leaders genuinely want a floor under the industry before a bad incident writes the law for them. The incentive problem does not care about sincerity at the signing table. It cares about what the next product review rewards.
A Whistleblower’s Memory Of Flexible Systems
Haugen’s credibility on this topic comes from a specific job. She worked on civic misinformation, saw internal research, and decided the public should see it. The documents, later dramatized in a film opening on October 9, described a company that understood certain harms and still struggled to change the machine that produced them. She later wrote about the personal cost of speaking. None of that makes her an oracle on model weights. It does make her a witness to how large platforms behave when internal knowledge and external promises diverge.
One cluster of that research concerned teens. Internal work suggested the product could worsen body image and mood for some young users. The public line was warmer than the internal picture. That gap is the pattern she is now mapping onto AI. Know the risk. Describe a safeguard. Leave the metric that creates the risk mostly intact.
You do not need to relitigate every old argument about feeds to use the lesson. The lesson is structural. When a system is optimized for time, clicks, or completed tasks, safety language has to fight the optimizer. If safety is a side memo, the optimizer wins. If safety changes what “good” means in the weekly review, the optimizer can be pointed somewhere else.
Why AI Talks Feel Different From The Last Cycle
Haugen was careful not to paint the current moment as a repeat. She said AI companies have engaged safety concerns earlier than social companies did. Some of that is genuine. Some of it is scar tissue. Executives watched hearings, leaks, and reputation damage. They would rather shape the rules than inherit them after a crisis.
There is also a technical difference. A feed can harm quietly for years. A model that writes exploit steps, imitates a trusted voice, or steers a medical decision can fail in public, fast. The blast radius is part of why the conversation started sooner. Fear is not a governance model, but it does change calendars.
Still, earlier engagement is not the same as durable oversight. A company can publish a safety page, join a pledge, and keep the evaluation set small enough that the bad case never appears in the slide. I have watched teams celebrate a red-team report that never left the building. The report existed. The product did not change.
Independent Auditors And The Angel Problem
Haugen singled out a push, associated with Anthropic and its chief executive Dario Amodei, for independent auditors. Her reading was blunt. If you are given room to cut corners, even the angels start cutting them. Outside eyes are not an insult. They are a design feature for humans under pressure.
If you are given the space to cut corners, even the angels among us begin to cut corners.
Frances Haugen, on why outside auditors matter
An independent audit is not a magic stamp. It can be captured, delayed, or scoped so tightly that it misses the product people actually use. Done well, it still changes the room. Someone who does not report to the launch owner asks to see the failure cases. Someone writes down what was refused. That record is awkward, which is the point.
Investors should care about this even if they never read a model card. Audit rights, incident reporting, and the power to pause a deployment are governance questions. They sit next to board oversight and insurance terms. A firm that invites a real auditor is telling you something about how it expects to be wrong. A firm that only invites a friendly reviewer is telling you something else.
Letter Versus Spirit In Everyday Product Choices
Abstract pledges become real in small meetings. A team wants to ship a companion feature that remembers personal details. The written rule says do not store sensitive data without consent. The designer adds a toggle. The toggle is buried. Most people leave the default on. Has the rule been met?
Or take youth access. A clause says the model should not target minors with harmful content. The product is rated for general audiences. Age checks are weak. The harmful pattern is not a single banned phrase. It is a slow drip of comparison, advice, and flattery. Letter met. Spirit missed. This is the fence problem in product clothing.
Another version shows up in enterprise deals. A customer wants the model to draft performance reviews. The safety policy forbids discriminatory output. The model still mirrors biased historical text unless someone measures that slice on purpose. Nobody prohibited the contract. Nobody built the test. The gap is operational, not rhetorical.
Spirit test, in plain language: What harm was the clause trying to stop? Where else can that harm appear? Who is paid to look there after launch? What happens if they find it?
If a company cannot answer those four lines, the signature is ahead of the system. That is not a scandal by itself. It is a timetable. The scandal starts when the timetable never moves.
What Markets Should Watch After The Signing
Public markets price stories quickly and operations slowly. A signing day can lift a narrative about responsible leadership. The bill arrives later, in incident costs, customer churn, regulatory follow-up, or a product delay that was actually a safety delay. Haugen’s interview is a reminder to look past the photograph.
A few signals are more useful than the press statement.
- Does the firm name an outside auditor with a scope that includes deployed products, not only lab models?
- Are safety incidents reported on a clock, with enough detail for customers to judge?
- Did any launch slip because a review failed, and was that slip explained?
- Are youth, health, and financial use cases held to a higher bar than novelty demos?
- Is the safety team measured on prevented harm, or on the speed of policy approvals?
None of those items require a leaked cache. They show up in filings, customer contracts, and the way executives answer a follow-up question. If every answer is a restatement of the pledge, you are still at the fence, not past it.
| Signal | Letter-only version | Spirit version |
| Policy text | Banned phrases listed | Harm patterns defined and retested |
| Audit | Internal review, summary only | Independent scope, published method |
| Youth risk | Age gate on paper | Measured outcomes after launch |
| Incident path | User can file a ticket | Clock, owner, and public lesson |
| Incentives | Launch date unchanged | Ship blocked when tests fail |
The table is a blunt tool. Real companies will sit in the middle. The useful habit is to ask which column the next decision lands in. Middle is fine for a quarter. Middle forever is how the last platform era aged.
Social Platforms Wrote The First Draft Of This Argument
It is tempting to treat AI as a clean break from social media. The talent overlaps. The advertising model still funds a large share of consumer tech. The habit of testing engagement before testing wellbeing did not retire when chat windows arrived. Haugen’s career sits on that bridge, which is why her skepticism travels.
Social companies learned to speak the language of wellbeing while the ranking systems kept their old job. Some changes were real. Harmful content policies grew. Transparency reports appeared. Teen tools were added after pressure. The core argument from critics remained that the business model and the safety model were negotiating, and the business model had the louder voice in ordinary weeks.
AI labs are not identical to ad-funded feeds. Several sell subscriptions or enterprise contracts. That can align a firm with a customer who wants fewer disasters. It can also align a firm with a customer who wants fewer refusals. Both pressures are already visible. A bank wants caution. A growth team wants the model to say yes. Self-regulation has to survive that argument on a random Thursday, not only on signing day.
The Film, The Book, And The Timing
The cultural timing is not accidental. A dramatization of the document leak opens on October 9, and Haugen has already told her version in a book about deciding to speak. Stories like that reset public memory. They remind people that internal knowledge and public language can diverge for a long time before anyone outside notices.
I would not treat a movie as evidence. I would treat it as a mood that executives cannot fully control. When audiences are freshly reminded of a gap between research and rhetoric, a new pledge gets read with less charity. That may be unfair to teams doing serious safety work. It is also predictable. Trust is a lagging indicator.
For the firms that signed, the practical response is not a sharper statement. It is a visible refusal. Ship one less feature. Publish one uncomfortable metric. Invite one auditor who can say no. Those moves are quieter than a premiere, and they age better.
Where Self-Regulation Has Worked Before
It would be lazy to claim voluntary rules never work. Aviation, some corners of finance, and parts of medical device practice built cultures where peers punish sloppiness because the downside is shared. The common ingredients are not slogans. They are licensure, insurance, incident databases, and a professional cost for hiding a near miss.
AI does not have that stack yet. Talent moves fast. Models are partly opaque. Harm can be diffuse, which makes blame easy to dodge. A pledge can be the start of a stack. It cannot substitute for the stack. Haugen’s “step up and comply” line is really a request to build the missing pieces while the signature is still fresh.
Recent policy research on emerging technology, read across several fields, keeps landing on the same boring conclusion. Voluntary codes change behavior when someone outside the firm can compare claims with practice. Without that comparison, codes mostly change vocabulary. I share that reading. Vocabulary is cheap. Comparison is not.
The Human Stakes Behind The Governance Talk
Governance language can sand the edges off real harm. A teen asking a model how to disappear weight. A worker whose review was drafted from biased patterns. An older adult talked into a payment by a voice that sounds like a grandchild. These are not edge cases in a footnote. They are the reason the fence was discussed at all.
Haugen’s earlier disclosures mattered because they connected a product metric to a person’s week. AI safety debates drift toward extinction scenarios and lab benchmarks. Both matter. So does the ordinary week. A self-regulatory deal that only speaks to frontier risk and ignores deployed, boring harm will repeat the miss she already documented.
Families do not experience a model as a research program. They experience it as an answer that arrived with confidence. Confidence without a trail is a product choice. Companies that want the spirit of a safety deal will make the trail visible, even when the answer is “we do not know” or “we will not do that.”
A Practical Reading For Boards And Buyers
If you sit on a board, buy enterprise tools, or allocate capital, the interview is a checklist in disguise. Do not ask whether the firm signed. Ask what would embarrass the firm if an outsider looked next month. Then ask whether anyone is paid to look.
Buyers can write the spirit into contracts without waiting for a statute. Require evaluation on the customer’s own risky tasks. Require notice when a model update changes refusal behavior. Require a human path when the system is used on health, credit, hiring, or minors. Those clauses are dull. Dull clauses are how pledges become operations.
Boards can ask a harder question. If we removed the safety team’s ability to delay a launch, what would ship this quarter? If the answer is “almost everything on the roadmap,” the team is advisory. Advisory teams do not hold a fence line. They decorate it.
Board question: Can safety delay a launch without a founder override? If not, the pledge is a brochure.
Founder overrides are sometimes correct. A safety team can be wrong, slow, or captured by fear. The point is not veto theater. The point is a recorded reason when the override happens. Silence plus speed is how corners get cut by people who still think of themselves as careful.
What “More Proactive” Should Mean Next Year
Haugen’s concession that AI firms have been more proactive is a door, not a medal. Next year will show whether the door was used. Proactive, in a way that would satisfy the spirit test, would look roughly like this.
- Public incident summaries that name the product, not only the lab.
- Third-party audits with a scope the auditor helped write.
- Youth and health evaluations that survive contact with real users.
- A habit of killing features that pass a demo and fail a week of use.
- Clear limits on memory, imitation, and persuasion, written for customers.
Miss most of that list and the proactive era was a communications cycle. Hit most of it and the industry will have done something social platforms mostly postponed. Either outcome is legible. You will not need a leak to tell them apart. You will need patience and a memory for what was promised in the signing week.
I keep coming back to the angels line because it is the least cynical sentence in the interview. It assumes good intent and still demands a structure. That is a more adult standard than “trust us” or “ban it.” Good intent plus a gap in the fence is how ordinary harm scales. Structure is how intent survives a bad quarter.
The Risk Of Writing Rules For Yesterday’s Model
There is another trap. Agreements freeze a picture of the technology at signing. Agents, memory, and tool use are moving faster than the nouns in most policy drafts. A rule aimed at a chatbot can miss an agent that takes actions. A rule aimed at generated text can miss a voice that calls your parent. Spirit compliance means updating the tests when the product shape changes, not waiting for the next summit.
This is where flexible systems punish static documents. The document cannot name every path. The company can still own the duty to look for new paths. If “not expressly prohibited” becomes the operating slogan, the agreement will age into a curiosity. Haugen basically dared the signers not to adopt that slogan.
Regulators will read the same dare. Voluntary deals buy time. They do not buy indefinite patience. If a visible harm lands in a domain the pledge claimed to cover, the political cost of self-regulation rises overnight. Firms that treated the text as a ceiling will wish they had treated it as a floor.
A Note On Fairness To The People Building The Systems
It is easy to write as if every lab is a single will. They are not. Safety researchers inside these companies argue, lose, win, and sometimes leave. Shipping teams are not a monolith of indifference. Some of the most detailed risk work I have seen came from people who stayed. Painting them as props in a pledge photo is inaccurate and, worse, useless.
The fairer claim is about structure. Individuals can care and still lose to a launch process that never gives them a real vote. Haugen’s story is partly about that loss. The auditor argument is an attempt to give internal dissent a witness. Without a witness, dissent becomes a Slack thread. With a witness, dissent can become a condition of release.
Credit belongs where tests are hard and results are shared. Skepticism belongs where every result is a victory slide. Both can be true in the same firm, in the same month. Readers should resist the urge to pick a permanent hero. Watch the next release instead.
How Readers Can Pressure The Spirit, Quietly
Most people will not audit a lab. They can still refuse the fog. Ask the product what it will not do. Notice whether the answer is specific. Turn off memory you do not need. Keep health and money decisions in a lane where a human remains accountable. Teach teenagers that a fluent reply is not a cared-for reply.
Those habits do not replace company duty. They stop you from outsourcing judgment to a system that has not earned it. Haugen’s larger point is that companies asked for trust by signing. Trust is a response to behavior over time, not a coupon attached to a ceremony.
If you work adjacent to these tools, you have a sharper lever. Write the evaluation your vendor hoped you would skip. Send the failure case back with a date. Procurement language is unglamorous, and it moves more product behavior than a panel discussion. I have found that a single rejected renewal does more than a dozen principles documents.
What Would Count As Stepping Up
Stepping up is a phrase that dies if it stays vague. Here is a concrete version, grounded in the concerns Haugen raised rather than in a fantasy statute.
First, map the harms the agreement implies, including the ones next door to the written bans. Second, test those harms on the products people use, not only on a benchmark with a friendly name. Third, let an outsider see the method. Fourth, tie a real decision, delay or redesign, to a failed test. Fifth, say so in public without burying the lesson in a footnote.
Do that, and the fence reaches the tree line. Skip it, and the signature remains a photograph. The technology will keep being flexible. The only open question is whether the institutions around it will be.
A promise that cannot survive an unflattering test was never a safety system. It was a schedule.
Haugen did not claim the signers will fail. She claimed the history of flexible products gives them a well-worn path to technical compliance and practical evasion. The next few product cycles will show which path they take. Markets, parents, and the people building the models all have a stake in the answer, which is why a Friday interview about a pledge deserves more than a news cycle.
Watch the gap beside the fence. That is where this story will actually be decided.
Getting rich is easy. Stay there, that's difficult.
September Payrolls Miss: Can Bitcoin Price Rally?