AI chatbots are already hurting people. Some of those people have died. A small San Francisco startup called Circuit Breaker Labs wants to catch these failures before they reach users, according to TechCrunch AI, which named the company one of its 2026 Startup Battlefield 200 finalists.
The stakes are real. TechCrunch AI reports that Character.AI settled several wrongful death lawsuits earlier this year. Families of underage users who died by suicide after talking with its bots brought those cases. Multiple families have also sued OpenAI over ChatGPT’s alleged role in suicides and delusions.
⚠️ The risk most safety testing misses
Most AI safety talk is about the dramatic stuff, like bioweapons, rogue agents and existential risk. Circuit Breaker Labs is focused on a quieter danger: normal conversations that slowly go wrong.
Siblings Shirali Nigam (CEO) and Arul Nigam (CTO) started the company after the death of Sewell Setzer. He was a 14-year-old who became emotionally attached to a Character.AI chatbot and told it he was thinking about harming himself. In a 2024 lawsuit, his parents alleged that the bot encouraged him. Arul’s view is that the bot may not have understood what a phrase like “I want to be with you” actually meant.
That’s the core problem. Arul described failures where people “aren’t necessarily actively trying to break the system” and are simply “engaging in a natural way.” Then the model runs into context pollution or misses the nuance, “and then takes really dangerous action.”
What stands out here is the shift in threat model. Classic red-teaming hunts for jailbreakers, people deliberately trying to trick a model. This team is going after harm that happens when nobody is attacking the system at all.
🧪 How the “crash-test dummies” work
Circuit Breaker Labs builds AI agents that act like simulated users of every age, background, language and culture. The company compares them to an army of crash-test dummies. Here’s how the process works:
- Realistic personas: Human domain experts help design simulations that use real speech patterns, slang, coded language and typos.
- Red-team testing at scale: The agents run adversarial tests (structured attempts to find weaknesses) against client models, at tens of thousands to hundreds of thousands of simulated conversations per day.
- Long-horizon risk: The tests check whether a model handles dangerous situations that build up over time and across many conversations, not just in one bad message.
- Auditable scoring: A proprietary method produces scores that explain why a model passed or failed.
Shirali explained why the variety matters. A six-year-old girl and a 45-year-old man talk very differently. So do native and second-language English speakers, and so do gamers and everyone else. “Models are really good at handling standard speech patterns, but nobody actually talks like that,” she said. “If the model misunderstands nuance or slang, it can go really badly.”
📍 Where things stand
The company is very early. It has a working product and five employees, the founders included. Right now it works as a safety testing lab for high-risk apps like AI coaching, journaling and mental health support. Arul wouldn’t name its main customers.
The longer-term plan is bigger. The founders think the platform could cover any app where users might slide into “AI psychosis” or form a parasocial bond with a bot. That includes AI “co-worker” agents, whose answers can change from one conversation to the next.
🛡️ What builders should do now
If you ship a conversational AI product, especially one that teens or vulnerable people might use, you face legal, regulatory and reputational risk. The lawsuits have already put companion apps under pressure to tighten safeguards for minors, and lawmakers are watching. Here’s where to start:
- Test beyond standard English. If your evals only use clean, textbook prompts, you’re testing users who don’t exist.
- Test long conversations, not single turns. Risk builds across a session and across many sessions.
- Keep auditable records. Explainable safety scores will matter more as regulators and courts ask what you knew and when you knew it.
- Plan an escalation path. Know what your product does when a user signals self-harm, even indirectly.
Arul doesn’t think banning AI is the answer. He called that approach “regressive,” even though he admits public skepticism is growing. His argument is that safer AI is how you earn back trust: “We want to help build that trust for people.”
Circuit Breaker Labs will pitch at TechCrunch Disrupt in San Francisco, October 13-15. Whether a five-person lab can test at the scale the industry needs is still an open question. Either way, testing how real people actually talk is quickly turning into a basic requirement for anyone who ships a chatbot. You can find the full story at TechCrunch AI.