Build real pipelines on a real lakehouse. Explore the data, write the SQL, re-run it like a scheduler would, and get graded check by check, not by multiple choice.
Daily batches of order updates land in orders. Keep order_facts holding the latest version of each order, without refunds.
Your script may run twice for the same batch: it must be idempotent.
def validate(ws): ws.load("orders", "day_1") ws.run_submission() ws.load("orders", "day_2", mode="append") ws.run_submission() ws.run_submission() # idempotent? ws.check("latest version", ws["order_facts"].matches_reference(key="order_id"))
On the engines you'll use at work, over an Iceberg lakehouse
Not toy puzzles. Daily batches, late data, re-runs and dimension tables, graded the way your team would review them.
Batches land, your script runs, more data lands, it runs again. Problems test incremental loads, MERGE, deduplication and idempotency, not just one SELECT.
See which checks pass, why the others fail, and your score. Hidden cases never leak expected rows.
Read-only scratch tabs on the sample. Run a selection, save a result as a table, chain the next query on it.
Typed columns, row numbers, sorting, copy, and a panel you can drag open like a terminal.
Every run is sandboxed and every query is checked by policy and catalog grants. Your code never reaches hidden data.
One-click SQL formatting, ⌘K search, and a sample workspace that's ready in about a second.
Numbers within a tolerance, strict types, keys named in the feedback. No more failing on 0.0005 rounding.
The loop you'd follow at work, in one browser tab.
Open a scratch tab, look at each batch, run a selection of your query.
Write the pipeline: CREATE, MERGE, INSERT. Run it on the sample and inspect every table it leaves.
Hidden cases replay the scenario. Every check is graded; partial credit shows what's left.
The same scenarios that train one engineer scale to a whole cohort, a new team, or a fair interview loop.
Give every student a sandboxed lakehouse. Assign scenarios, set due dates, and see who's stuck, without grading SQL by hand.
Turn your real tables and past incidents into practice problems. New hires learn your data shapes before they touch production.
Every candidate gets the same scenario and the same checks. Compare scores and see exactly which requirements each one met.
Practice for free. Upgrade when you want speed, history and the newest engines.
Everything you need to start practicing.
For engineers preparing for the next role.
Author private problems on your own data.
Teach on a real lakehouse instead of slides. Teachers and teaching assistants are always free, your first class is on us, and students pay a fraction of Pro, billed to the school per term.
Teachers & TAs free · cancel any term
Yes. Every published problem can be solved and graded on the Free plan. Pro adds priority, parallel runs and history.
Trino SQL (one SELECT, or a full script with CREATE, MERGE and INSERT) or PySpark (return a DataFrame from solve(spark), or write the tables yourself). Both are graded by the same checks; dbt and Airflow are next.
No. Runs are sandboxed, every query passes a policy check, and hidden data and expected tables live in a catalog users can't read.
On the Team plan you author private problems: upload sample batches, write a reference solution and a short validate.py.
Teachers are free. Create private problems or reuse the catalog, assign them with a due date, and follow every student's checks. Students get everything in Pro.
Yes. Every candidate gets the same scenario and the same hidden checks, and you get a per-check report for each one.
Pick a scenario, explore the sample, and ship something that would survive a re-run.
Start practicing, free