A newly filed lawsuit alleges that xAI, the artificial intelligence company founded by Elon Musk, trained its Grok models on child sexual abuse materials (CSAM). The complaint, submitted on a Wednesday, arrives as regulators and courts examine the scope of the problem, and after some Grok users have been arrested in connection with abusive imagery.
The plaintiff, identified only as Jane Doe, states that she was of preschool age in the early 2000s when adult men repeatedly assaulted her to produce CSAM sold to offenders online. In the years since, her images have been catalogued and hashed by organizations including the National Center for Missing and Exploited Children (NCMEC) and the Canadian Centre for Child Protection (CCCP).
To protect herself, Doe elected to receive alerts through the US Department of Justice Victim Notification System whenever she may be identified as a victim in a new criminal investigation. According to the complaint, she has received numerous such alerts over the years. She was nonetheless alarmed when the CCCP informed her that it had identified AI-generated CSAM on xAI that depicted her.
Claims About Stored Outputs and Continuous Training
The complaint alleges that messages were discovered on online forums “between offenders chatting about creating AI generated CSAM of Plaintiff and other similarly situated known, legacy, victims of CSAM.” Doe contends that xAI has not only made it easier to generate additional violative images from the most traumatic period of her life, but that the company also allegedly stores the images Grok produces and uses those outputs to further train the model.
As a result, she believes Grok has been trained on both the original images that have followed her for more than 20 years and the newer AI-generated versions. This is described as the first case to accuse xAI of training on CSAM. The complaint does not provide extensive detail on that specific allegation, stating only that “CSAM depicting Plaintiff with its longstanding well-known hash values has been used as a part of the dataset used by xAI.” A press release from Doe’s attorneys similarly asserts that her images appeared on a CSAM Hash List maintained by NCMEC and that “that same material” was allegedly “part of the dataset xAI used to build Grok’s image and video generating capabilities.”
How Public Posts May Feed the Model
The filing offers more detail on the claim that Grok trains on AI-generated CSAM. “Because Grok’s terms treat public X posts and Grok’s own outputs as training data by default, publicly posting an image does not just expose it to viewers, but also feeds [it] directly into the pipeline xAI uses to train and improve its model and thereby generate further images,” the complaint states.
The lawsuit notes that while xAI filters violent content out of Grok outputs to keep it from training data, the company’s terms do not specify whether CSAM, non-consensual intimate imagery (NCII), or NSFW material are treated as “excluded categories.” The complaint argues that fully removing a training example’s influence from an already-trained model is technically difficult and something xAI has not publicly claimed to have done, so any CSAM ingested before takedown likely continued to shape the model’s outputs even after the original images were removed from public view.
Source
Image: arstechnica.com