The next battleground in AI isn’t a bigger model. It’s your company’s Slack history. According to The Information, OpenAI and Anthropic’s hunger for high-quality, real-world data has turned the messy internal chatter of startups into a genuinely valuable asset. Private conversations that used to sit forgotten in a channel are now something the biggest labs will pay attention to.
What stands out here is the shift in what counts as valuable data. The open web has largely been scraped dry. Blog posts, Reddit threads, code repos, and public documentation are already baked into today’s models. What the labs are short on is the stuff humans never publish: how real teams argue through a decision, debug a problem, hand off work, and explain their reasoning in plain language. Slack threads capture exactly that.
📊 Why this data is suddenly worth so much
The frontier labs are running into a wall that money alone can’t knock down. They’ve used most of the good public text, and synthetic data (models training on model output) tends to get stale or degrade over time. Human-generated workplace conversation fills a specific gap.
- It’s authentic. Real people solving real problems, not polished marketing copy.
- It’s reasoning-rich. Threads show the back-and-forth, not just the final answer.
- It’s domain-specific. A biotech startup’s Slack teaches the model things no textbook covers.
- It’s scarce. This data is locked inside private workspaces, so whoever controls it holds leverage.
That scarcity is what turns a startup’s archive into a bargaining chip.
⚖️ The tension nobody has solved yet
Here’s where it gets uncomfortable. That Slack history isn’t just the company’s. It’s full of employees’ words, customer details, and third-party information shared in confidence. Selling or licensing it for AI training raises real questions about consent and privacy that most startups haven’t thought through.
The Information’s reporting points to a market forming faster than the rules around it. Regulators in the EU and several US states are already circling how training data gets sourced. A startup that hands over its internal communications today could be handing over data it never had clean rights to license in the first place.
🧭 What this means for the next two years
Expect data to become a line item on the startup balance sheet. Proprietary conversation logs, support tickets, and internal wikis will get valued the way patents and code do now. A few things are likely to follow:
- Data brokers move in. Middlemen will spring up to package and resell workspace data to labs, the same way they did for web scraping.
- Contracts change. Employment agreements and customer terms will start spelling out AI training rights explicitly.
- Platforms fight back. Slack, Notion, and similar tools will clarify who owns what, and may start monetizing the data themselves.
This is significant because it flips the usual story. For years the concern was labs quietly taking data. Now startups are being courted to give it up willingly, for cash.
✅ Practical takeaways
If you run or work at a company sitting on years of internal data, treat it seriously before anyone comes knocking.
- Audit what you have. Know what lives in your Slack, your docs, and your support logs.
- Check your rights. Review employee and customer agreements before you consider licensing anything.
- Don’t undersell. If this data is genuinely scarce, it has real leverage. Price it like an asset, not a giveaway.
- Weigh the reputational cost. Employees and customers may not love learning their words trained a commercial model.
The labs need what you already own. That’s a stronger negotiating position than most founders realize, and it won’t last forever. The window is open now, while public data runs thin and synthetic alternatives fall short. For the full picture, the original reporting from The Information is worth your time.