63% of Free AI Tools Hide User Data Training, Cybernews Finds

63% of Free AI Tools Hide User Data Training, Cybernews Finds
Key Takeaways

  • The Cybernews AI Trustworthiness Ranking 2026 found that nearly two-thirds of 500 AI companies surveyed do not clearly disclose whether user input is used to train their models.
  • Free AI tiers, including consumer ChatGPT, may use submitted content for model training by default; business tiers typically do not, a distinction providers rarely surface at sign-up.
  • A July 2026 Singapore PDPC advisory confirmed that data originally published for other purposes, including forum posts and professional profiles, can be legally incorporated into AI training sets without the subject’s knowledge.

When a Canadian professional services firm pasted a draft severance letter, real names and real dollar amounts, into a free AI chatbot, that data likely entered the model’s training set permanently. A report published August 18 by Cybernews found this kind of silent data transfer is the norm: its AI Trustworthiness Ranking 2026 found that 42% of 500 AI companies surveyed made no reference to training on user data in their privacy policies at all, and another 21% gave only vague explanations.

The Real Price of Free AI

Mike Pearlstein, CEO of Fusion Computing, wrote in an August 10, 2026, blog post that the price tag is zero but the trade is your business data. The tier a user is on, free or paid, determines whether submitted content stays private or feeds the model. Consumer ChatGPT may use submitted content to train its models; business tiers typically do not by default. That distinction is rarely surfaced during sign-up, according to the Cybernews report.

Once In, It Stays In

Data absorbed into a large language model does not behave like a database record. It becomes distributed across billions of parameters, with no individual entry to locate and no clean removal path. As Pearlstein put it in his August 10 post, once content enters the training corpus, it cannot be pulled back. Retraining from scratch to excise a specific input is technically possible but practically unrealistic for most providers.

The practical risk is concrete: an unpatented product concept, a client’s financial details, or a confidential negotiation submitted through a free tool that trains on user data could surface in future model outputs, even indirectly. De-identification offers limited protection, since techniques exist to infer original data from model behaviour. For more on how AI training data intersects with legal liability, the regulatory picture is still forming.

More Than Your Chats

Direct user input is only one exposure point. Many models are trained on web-scraped data gathered before a user ever types a prompt. Singapore’s Personal Data Protection Commission published advisory guidelines on July 20, 2026, clarifying how its Personal Data Protection Act applies to AI training, including web scraping and the reuse of data originally provided for other purposes. Publicly available data accessed without restrictions can generally be used; data behind paywalls requires additional assessment. The guidelines acknowledge that a significant portion of what enters AI training was never explicitly offered for that purpose by the people it concerns.

Content published online, forum posts, professional profiles, published work, may already be part of training sets the subject had no say in. That exposure predates a user’s first interaction with any tool. The question of who owns that data and who bears liability is actively contested in courts and legislatures.

Surveillance Pricing

Data exposure through AI tools runs alongside a separate but related risk: AI-driven personalised pricing. At a Senate Judiciary Committee hearing in 2026 titled “Your Data, Their Profit: The Consumer Cost of AI Surveillance Pricing,” witnesses described how companies use AI to set prices based on individual consumer profiles. Lindsay Owens, President and CEO of the Groundwork Collaborative, testified that Instacart offered different prices for identical groceries to different shoppers simultaneously, with potential annual differences of around $1,200 for an average family. Target’s shopping app was cited at the same hearing as tracking real-time location and charging customers as much as $150 more for the same product if they were already in the store’s parking lot.

The Federal Trade Commission issued a draft policy statement on August 19, 2026, targeting AI-driven personalised pricing. It would require firms to disclose when prices are personalised, what data drives that personalisation and what data types are involved. The FTC’s position is that consumers expect a listed price to be a price, not an estimate of their willingness to pay based on browsing history.

Regulations Lag the Practice

California’s AI Transparency Act and the EU AI Act’s Article 50 transparency rules do not directly target the internal data practices the Cybernews report flagged. Regulations are beginning to address what AI produces; what goes into making it, user data, scraped content, implied consent, remains inconsistently covered. The Cybernews finding that nearly two-thirds of surveyed AI companies do not clearly disclose their training data practices sits in exactly that gap.

Practical Steps

Pearlstein’s August 10 post recommended reading the privacy policy before pasting anything sensitive into a free tool, specifically looking for language about whether user input is used for model training or fine-tuning. Vague or absent language on that point is itself an answer.

For confidential material, client data, financial figures or unpublished work, the practical options are a paid enterprise tier with explicit data protections, a locally hosted model, or not using an AI tool for that task. The cost difference between a free tier and a business subscription exists partly because the business subscription is not subsidised by user data. Teams evaluating where AI fits in their data governance stack are increasingly treating tier selection as a data risk decision, not just a budget one.

Alex Chen
Alex Chen

Alex covers AI tools, apps, and consumer technology for Auton AI News. With a focus on making AI accessible, Alex helps everyday readers understand and use the latest AI developments.

📰 Journalists welcome — cite Auton AI News with attribution. Press & Media → | press@autonainews.com