Monday, August 3, 2026

AI & Models

OpenAI reportedly asks contractors to upload real work samples

OpenAI and Handshake AI are reportedly asking contractors to upload real-world work samples to train AI models, raising significant intellectual property and confidentiality concerns.

OpenAI reportedly asks contractors to upload real work samples

The push for high-quality training data has led artificial intelligence companies to seek increasingly specific real-world inputs. OpenAI and training data company Handshake AI are reportedly asking third-party contractors to upload real work that they did in past and current jobs. According to a report in Wired, this practice appears to be part of a broader strategy across AI companies that are hiring contractors to generate high-quality training data in the hopes that this will eventually allow their models to automate more white-collar work.

In OpenAI’s case, an internal company presentation reportedly asks contractors to describe tasks they have performed at other jobs and upload examples of real, on-the-job work that they have actually completed. The presentation specifies that these examples should include concrete outputs rather than summaries of the files, requesting the actual files such as Word documents, PDFs, PowerPoint presentations, Excel spreadsheets, images, or code repositories. To address privacy concerns, OpenAI reportedly instructs contractors to delete confidential and personally identifiable information before uploading. The company points them to a ChatGPT tool called “Superstar Scrubbing,” which is designed to remove personally identifiable information (PII) from data before submission.

However, this data collection method introduces severe legal vulnerabilities. Intellectual property lawyer Evan Brown told Wired that any AI lab taking this approach is “putting itself at great risk.” Brown noted that the strategy requires a significant amount of trust in contractors to decide what is and is not confidential. By relying on individual contractors to filter out proprietary data, AI companies may inadvertently ingest protected corporate intellectual property, exposing themselves to potential litigation. An OpenAI spokesperson declined to comment on the report.

Why it matters

This practice highlights the aggressive push by AI companies to secure high-quality training data, while simultaneously exposing them to significant legal and confidentiality liabilities regarding intellectual property.