
The deterministic layer for frontier intelligence
Despite massive investment in commercial AI, organizations often find that demonstrated value is elusive, primarily due to the non-deterministic risk inherent to generative models. CTGT is the deterministic governance layer that enables the most important global institutions to deploy AI workflows with confidence.
Born out of Stanford University research, we provide the control plane that makes it possible. A lightweight, model-agnostic system that enforces policy, prevents drift, and produces auditable decisions in real time. When benchmarked on HaluEval, the CTGT Policy Engine (paired with GPT-120B OSS) outperformed frontier models (Gemini 3 Pro Preview, Claude 4.5 Opus and 4.5 Sonnet) at drastically lower compute cost.
While we sit on the edge of AI research, CTGT brings frontier intelligence into real-world environments. We apply cutting-edge theory directly in production to make large language models more reliable, controllable, and performant in practice.
Our mission is to bring models to the level of performance and accountability required by the Fortune 500. By bridging the gap between LLM capabilities and domain-specific requirements, we unlock the true potential of generative AI to solve the most pressing problems in our world today.
Not your average fixed-point internship.
Frontier models are now usually right and occasionally confidently wrong, and they cannot tell you which is which. A model that is 95% reliable is useless in the settings we serve, the same way a self-driving car that avoids most accidents is useless. CTGT's research function exists to close that gap. Our founding research stems from feature learning in neural networks, and we use that machinery to extract and steer features at runtime, on open and closed-weight models, without training a new artifact for every behavior.
As a research intern, you will own one hard problem inside this program from end to end. You will not be handed a labeling task or a notebook to babysit. You will take a real research question, like a better way to find what a model represents, intervene on it, or bound how wrong the system can be, design an approach, implement it against real models, and prove or disprove it with evidence that holds up. You will sit directly with the engineers building the Policy Engine, present in our weekly research review, and be expected to form opinions, ask hard questions, and take problems further than they were handed to you.
We hold interns to the standard of a calculation that has to be right, not a demo that usually works; in practice this means limited ground truth, unverifiable intermediate steps, and failure modes that hide in the tails.
World-Class Backing: You will join a venture-backed company with institutional investors including Google's Gradient Ventures, General Catalyst, and Y Combinator.
Real Impact: You will work directly on the core systems that determine how models perform in the wild. Your work ships into real, high-stakes environments where governance, auditability, and performance are non-negotiable.
Autonomy & Trust: We operate with a high degree of trust. You are expected to form strong technical opinions and execute on them.
Apply through Work at a Startup with the single most interesting thing you have built or proven, such as a paper, a repo, or a calculation, and a sentence or two on why it mattered. We read every submission, this artifact is the signal.