HomeCompaniesDatoric

The secure training data R&D engine for physical AI

Datoric develops custom datasets for voice models, robotics, and world models, treating research, collection, verification, and production as one continuous process. We work closely with frontier model teams to turn emerging limitations and model failures into testable data hypotheses, while running our own experiments ahead of customer demand. This allows us to operationalize validated methods into repeatable collection systems at scale. Data is collected through private invite-only applications separated by modality, customer, and trust level. Every submission remains linked to the contributor, device, task, session, consent, rights, and processing history that produced it, giving our internal QA and fraud models the context to detect problems that may appear legitimate in the finished file. Each collection reveals new failure cases and quality signals that improve the systems behind the next dataset. Once a collection method is validated, we scale it up as a solution to the model failure and it also becomes a reusable data recipe for future custom projects or independently collected, rights-cleared data products.
Active Founders
Nikhil Reddy
Nikhil Reddy
Founder/CEO
CEO @ Datoric. Prev Quant & SWE Intern. Math/Econ/CS @ UChicago; Bypassed Google OAuth 2.1 and QA systems on data annotation platforms to automate training tasks in 2023, then disclosed the loopholes to the affected platforms.
Jeffrey Lin
Jeffrey Lin
Founder/CTO
CTO @ Datoric. Prev AI/ML & SWE Intern. Math & CS & Robotics at NYU. Created a bounding box algorithm for more efficient labeling as Head of Data for the NYU Robotics Team.
Company Launches
Datoric | Trustworthy data for the next generation of models
See original launch post

Datoric provides custom, security-first training data for voice models, robotics, and world models. Our private, project-specific contributor systems have led to a dramatic improvement in quality for customers and helped us generate nearly seven figures in revenue over the last 30 days.

Our background

In 2023, we built large-scale automations that completed paid LLM training tasks across major crowd-work platforms by bypassing QA testers. We earned six figures in three months, and even helped fix some of the vulnerabilities that we found. From this experience, we learned just how easily training data can be compromised, and how it can lead to unsafe model behavior.

https://youtu.be/GXuyd2PfsAY

The problem

Labs are searching for the high-quality, authentic human data their models need, but quality and provenance become harder to control at scale, especially when relying on outside data providers. Poor-quality or compromised data can distort model behavior, while unqualified contributors, automated submissions, unclear instructions, and weak consent records can still slip through QA. At the same time, data providers hold concentrated stores of customer data, contributor identities, and payments that attract fraud and security threats.

What we built

Datoric manages the full collection process for multilingual speech, egocentric video, and action-conditioned data. We collect this data through more than 300,000 active contributors who record conversations, everyday tasks, and other project-specific activities through our private apps. Because Datoric has no public marketplace or open upload system, every collection can run inside its own isolated environment, keeping customer projects separate and limiting unauthorized access. Each recording also remains linked to the verified contributor who created it and the consent they provided. This blocks automated or manipulated submissions and protects quality throughout the collection process rather than leaving it to final QA.

Traction

We’ve seen a major uptick in quality compared to other providers, and in the last 30 days, Datoric has generated nearly seven figures in revenue.

Our ask

We would love introductions to teams building voice models, robots, or world models that need difficult, custom human-generated data.

Send us your hardest collection specification at founders@datoric.com, or visit datoric.com.

Datoric
Founded:2025
Batch:Summer 2026
Team Size:2
Status:
Active
Location:San Francisco
Primary Partner:Harshita Arora