HomeCompaniesAlkera AI

Reliable and safe data engineering and data science agents

Alkera is a data engineering/analysis/science agent that works in your IDE or CLI. If you’re tired of agents that produce false, or even dangerous, reports about your data, then Alkera is for you. Our agent is equipped with your data’s history and lineage, your team’s knowledge, and tailored plugins for data apps such as Snowflake and dbt. Simply install and connect your services to begin doing trusted agentic data work. We also offer VPC & on-prem deployment options for when data can't leave your boundary.
Active Founders
Rick Gao
Rick Gao
Founder
Founder at Alkera AI, Inc. Previously at Verition Fund Management where he built financial models and optimizations for treasury financing. Previous computational biology research at Yale's Gerstein Laboratory.
Tony Li
Tony Li
Founder
Founder at Alkera AI, Inc. Previously Agentic Software Engineer at ByteDance, where he built a Python framework for multi-agent AI workflows and deployed production bots used across company teams. Former Quantitative Developer at Carthage Capital. BS in Computer Science and Mathematics from Yale.
Andrew Tran
Andrew Tran
Founder
Founder at Alkera AI, Inc. Prev SWE intern at Ramp and HRT working closely with data teams. CS B.S/M.S from Yale. Prev ML research at Yale's Gerstein Lab developing biology reasoning models.
Company Launches
Alkera - The data agent that you can trust
See original launch post

TL;DR: Data teams adopting AI agents want speed, but one wrong step can quietly ripple through your dbt models, dashboards, and business decisions before anyone notices. Alkera is an AI data agent for data engineers/analysts/scientists or anyone who wants to gather data insights that actually understands your data stack (column-level lineage, graph-based planning, and a living knowledge base), so it ships data work you can trust. It's #1 on UC Berkeley's DataAgentBench, which tests complex, data-oriented tasks.

https://www.youtube.com/watch?v=fQEVpQPYNxg

Coding agents can write data code that runs. For data engineering, knowing it's correct is the hard part. This is knowing that a transformation didn't quietly redefine a column three models downstream and corrupt the revenue dashboard that the exec team checks every Monday. For data analysis and data science, beyond data access and lineage problems, coding agents struggle with highly parallel work spread across multiple machines. When interacting with remote computing nodes, they do so only via fragile chains of SSH commands. Long processing and training runs have you more as a babysitter than a scientist.

Alkera is a data agent that works natively across your entire data stack, with the context and guardrails coding agents lack. Every change it makes to a query, transformation, model, or notebook is backed by a real understanding of your pipeline.

It's #1 on UC Berkeley's DataAgentBench, which tests complex, data-oriented tasks. You reach it through an IDE extension (VSCode, Cursor), a CLI, or the browser — wherever you already work.

What makes it different:

  • Column-level lineage across your data stack. Trace a single row from Snowflake, through your dbt models, to your Hex dashboard. Before the agent changes anything, it maps the blast radius — every downstream table and dashboard the change would touch.
  • Graph-based planning with flexible computation. Plan complex data tasks with parallel agents from a quick analysis to days-long modeling experiments to run across your machine and remote CPU/GPU nodes.
  • A living knowledge base. Alkera learns your schema docs, trusted queries, and the tribal conventions that never make it into documentation — written by both agents and humans, with humans verifying what the agent adds.
  • Production protection. Fine-grained permissions over every tool, including bash and SQL. Block a seemingly innocent SQL query with a dangerous UPDATE hidden away in a subexpression. Enforce OAuth to connect your warehouses to maintain existing data permissions.

Try it

Download Alkera at https://alkera.ai or read the docs at https://docs.alkera.ai. For data engineering, point it at your stack and watch it trace a change through your pipeline, lighting up everything downstream. For data analysis and data science, have it spin up parallel research agents to explore your data. Reach out to contact@alkera.ai to see how we can tailor Alkera to your specific data apps/tools and computing environment.

Our team

  • Andrew Tran studied CS at Yale. At Yale, he trained LLMs to reason over multi-modal biological data. At Ramp and Hudson River Trading, he worked closely with data teams to build out ergonomic tooling for data engineers and researchers.
  • Rick Gao studied CS at Yale. He was previously at Verition Fund Management, where he worked with internal data teams and counterparties to develop ingestion pipelines and internal financial models.
  • Tony Li studied CS and Math at Yale. During his time at Carthage Capital and ByteDance, he worked with data engineering teams to build data pipelines, data analysis, and agentic tooling.

Our ask

If you:

  • perform data engineering work on production pipelines with tools like dbt and Apache Airflow
  • perform data analysis/data science work in Python/R and with tools like Jupyter notebooks
  • manage a data team – engineers, analysts, or scientists
  • want to enable non-data experts to create data insights

Check out https://alkera.ai and reach out at contact@alkera.ai. We’d love to showcase a demo, learn more about your data teams’ pain points, and share more about how Alkera can help supercharge them.

Alkera AI
Founded:2026
Batch:Summer 2026
Team Size:3
Status:
Active
Location:San Francisco
Primary Partner:Ankit Gupta