NewReal-life scenarios: incremental loads, MERGE, idempotent re-runs→

Practice data engineering
like it's production.

Build real pipelines on a real lakehouse. Explore the data, write the SQL, re-run it like a scheduler would, and get graded check by check, not by multiple choice.

No credit card·Graded in your browser·Trino SQL and PySpark today, dbt and Airflow next
Problems / Incremental order facts
Sample workspace
Medium·61% acceptance
Incremental order facts
TrinoMERGEIncremental loads

Daily batches of order updates land in orders. Keep order_facts holding the latest version of each order, without refunds.

Your script may run twice for the same batch: it must be idempotent.

Schema · orders
order_idbigint
customer_idvarchar
amountdouble
updated_attimestamp(6)
Tables
⌕ Search tables…
Sources
▾orders3 rows
order_id
customer_id
amount
Datasets · raw batches
▸day_1
▸day_2
Saved results
▸_r1
4 tables · sample data
query.sql GRADEDscratch 1
▸ RunSubmit
1-- runs after every batch, may run twice
2CREATE TABLE IF NOT EXISTS order_facts (
3 order_id bigint, amount double, updated_at timestamp(6)
4);
5MERGE INTO order_facts t
6USING (
7 SELECT *, row_number() OVER (
8 PARTITION BY order_id ORDER BY updated_at DESC
9 ) AS rn FROM orders
10) s ON t.order_id = s.order_id AND s.rn = 1
11WHEN MATCHED THEN UPDATE SET amount = s.amount
12WHEN NOT MATCHED THEN INSERT ...;
Test resultsOutput 4
#123 order_idABC customer1.2 amount◷ updated_at
11a10.002026-03-01 09:00
22a25.002026-03-02 09:00
33b5.002026-03-01 11:00
44c40.002026-03-02 10:00

On the engines you'll use at work, over an Iceberg lakehouse

Trinograded todayPySparkgraded todaydbtsoonAirflowsoon
Product

Everything a real pipeline throws at you

Not toy puzzles. Daily batches, late data, re-runs and dimension tables, graded the way your team would review them.

day_1.csvday_2.csv append
ordersyour script ↻ ×2order_facts ✓

Real-life scenarios

Batches land, your script runs, more data lands, it runs again. Problems test incremental loads, MERGE, deduplication and idempotency, not just one SELECT.

Overlapping days50%
day 1: latest version of each order
day 2: latest version of each order×289 duplicated keys
one row per order89 duplicated values of order_id
raw orders are untouched
no refunds

Graded check by check

See which checks pass, why the others fail, and your score. Hidden cases never leak expected rows.

query.sql batches × ×selection+
SELECT customer_id, sum(amount) FROM day_2 GROUP BY 1
read-only · sample0.09s
c40
a25
b-5

Explore before you build

Read-only scratch tabs on the sample. Run a selection, save a result as a table, chain the next query on it.

123 order_id▲ABC customer_id▲1.2 amount▲
11a10.00
22a25.00
33b5.00
44c40.00
4 rows · 3 columns Copy

A results grid that feels like a warehouse

Typed columns, row numbers, sorting, copy, and a panel you can drag open like a terminal.

Sandbox · no network
Policy per query (OPA)
Catalog grants per user
Hidden test data: Access Denied

Isolated by design

Every run is sandboxed and every query is checked by policy and catalog grants. Your code never reaches hidden data.

⌕ incr ord ⌘K
Incremental order factsMedium

Formatted, searchable, fast

One-click SQL formatting, ⌘K search, and a sample workspace that's ready in about a second.

1.2 10.5000000001 ≈ 10.5 ✓◷ day: timestamp ≠ date ✗123 2 missing keys ✗

Every type, honestly compared

Numbers within a tolerance, strict types, keys named in the feedback. No more failing on 0.0005 rounding.

How it works

From first look to graded pipeline

The loop you'd follow at work, in one browser tab.

01

Explore

Open a scratch tab, look at each batch, run a selection of your query.

02

Build

Write the pipeline: CREATE, MERGE, INSERT. Run it on the sample and inspect every table it leaves.

03

Submit

Hidden cases replay the scenario. Every check is graded; partial credit shows what's left.

Use cases

Built for classrooms, data teams and hiring

The same scenarios that train one engineer scale to a whole cohort, a new team, or a fair interview loop.

Data Eng · Fall cohortdue Fri
Assignment: Incremental order facts
AM100%
KO80%
LS60%
JB35%
Schools & bootcamps

Teach data engineering on a real stack

Give every student a sandboxed lakehouse. Assign scenarios, set due dates, and see who's stuck, without grading SQL by hand.

  • Assignments with due dates and class progress
  • Auto-graded, check by check, with partial credit
  • Your own course problems, private to your class
Onboarding pathprivate · your data
✓Week 1Explore our lakehouse: orders, sessions
Week 2Rebuild the daily revenue job
Week 3Replay last quarter's late-data incident
Data teams

Onboard newcomers on your own pipelines

Turn your real tables and past incidents into practice problems. New hires learn your data shapes before they touch production.

  • Private problems built from your sample data
  • Replay incidents: late data, duplicates, re-runs
  • Track progress across the team
C7Candidate C-7Senior data engineer · 42 min83%
Handles late-arriving rows
Idempotent re-runs
SCD2 validity ranges
Hiring

Interview with real work, graded fairly

Every candidate gets the same scenario and the same checks. Compare scores and see exactly which requirements each one met.

  • Identical scenarios, objective grading
  • A per-check report for every candidate
  • Hidden tests stay hidden
For teams

Turn your real pipelines into problems.

Upload sample batches, write a reference solution, and describe the scenario in a few lines of Python. Shuffleet runs your reference to build the expected tables, so you never hand-compute answers.

  • ✓One entry point: def validate(ws), with helpers for the checks you need
  • ✓Custom checks in plain SQL: return the rows that break the rule
  • ✓validate.py runs sandboxed with no network; the reference is graded against itself
  • ✓Interview, onboard and level up your team on your own data shapes
See the Team plan
validate.pyvalidated
def validate(ws):
  ws.load("orders", "day_1")
  ws.run_submission()
  ws.load("orders", "day_2", mode="append")
  ws.run_submission()
  ws.run_submission()  # idempotent?
  ws.check("latest version",
    ws["order_facts"].matches_reference(key="order_id"))
Helpers
ws.load
ws.run_submission
matches_reference
unique
references
ws.check_sql
Validated · 5/5 checks on every case

Create a problem in four steps

About fifteen minutes from your data to a published scenario.
Source tables
orders
order_idbigint
amountdouble
Datasets → deliverables
day_1day_2→order_facts
1Describe the scenarioSource tables, the batches that fill them, the tables to build.
Caseday_1day_2
Sample3 rows3 rows
Overlapping days59 rows61 rows
Late refunds120 rowsUpload
2Upload the batchesOne CSV per test case and batch. The sample is visible, the rest stay hidden.
validate.pyvalidated
def validate(ws):
  ws.load("orders", "day_1")
  ws.run_submission()
  ws.load("orders", "day_2", mode="append")
  ws.run_submission()
  ws.run_submission()  # idempotent?
  ws.check("latest version",
    ws["order_facts"].matches_reference(key="order_id"))
3Write the reference and checksYour solution, plus validate.py: load, run, check, in order.
Validation #367validated
Sample5/5
Overlapping days5/5
Late refunds5/5
✓ Ready 9/9Publish
4Validate and publishYour reference runs through every case twice. All green? Publish.
Pricing

Simple plans, start free

Practice for free. Upgrade when you want speed, history and the newest engines.

Free

Everything you need to start practicing.

$0forever
Start for free
  • All published problems, in SQL and PySpark
  • Run, Submit and per-check feedback
  • Scratch queries on the sample
  • One run at a time
Most popular
Pro

For engineers preparing for the next role.

$15per month, billed yearly
Go Pro
  • Everything in Free
  • Priority queue: your runs start first
  • Up to 3 runs at once
  • Full submission history and scores
  • New engines first: dbt, Airflow
Team

Author private problems on your own data.

$32per seat / month, billed yearly
Start with your team
  • Everything in Pro
  • Private problems with validate.py
  • Team workspace and leaderboard
  • Admin roles and seat management
  • SSO (coming soon)
Schools & bootcampsFree for teachers

Shuffleet for Education

Teach on a real lakehouse instead of slides. Teachers and teaching assistants are always free, your first class is on us, and students pay a fraction of Pro, billed to the school per term.

  • Unlimited private course problems
  • Assignments with due dates
  • Class progress and leaderboards
  • Everything in Pro for every student
  • Grade export for your records
  • Campus SSO (coming soon)
Per student
$3per month, billed per school year
First class of up to 30 students free for a full term
Set up my class

Teachers & TAs free · cancel any term

Questions

Is it really free?+

Yes. Every published problem can be solved and graded on the Free plan. Pro adds priority, parallel runs and history.

What do I write?+

Trino SQL (one SELECT, or a full script with CREATE, MERGE and INSERT) or PySpark (return a DataFrame from solve(spark), or write the tables yourself). Both are graded by the same checks; dbt and Airflow are next.

Can my code see the hidden tests?+

No. Runs are sandboxed, every query passes a policy check, and hidden data and expected tables live in a catalog users can't read.

Can we use our own data?+

On the Team plan you author private problems: upload sample batches, write a reference solution and a short validate.py.

How does it work for a class?+

Teachers are free. Create private problems or reuse the catalog, assign them with a due date, and follow every student's checks. Students get everything in Pro.

Can we use it for interviews?+

Yes. Every candidate gets the same scenario and the same hidden checks, and you get a per-check report for each one.

Your next pipeline, graded in minutes.

Pick a scenario, explore the sample, and ship something that would survive a re-run.

Start practicing, free