For everyone who works with data

Practice the questions
data teams actually ask.

SQL, Python, PySpark, Databricks, Snowflake and Fabric — for data engineers, analytics engineers, data scientists, BI developers and platform engineers. Real schemas, hidden datasets, and a verdict that reports how your solution behaves at scale, not just whether it returned the right rows.

Deduplicate customers · PySpark · 50M row fixture
Accepted7 / 7 tests
Runtime1.42 s · beats 87%
Peak memory612 MB
Shuffle stages2
Output partitions8
Rows processed50,000,000
Your window sorts every row inside each partition. On skewed keys that sort dominates runtime — aggregate first and join back, or salt the key.
27Published questions
11Certification exams
9Technologies
50M rowsLargest hidden fixture

What you can practise

Counted from what is actually published, not a list of logos.

SQLSQL10 questions · 2 examsPythonPython9 questions · 1 examPySparkPySpark4 questions · 1 examDatabricks2 questions · 2 examsSpark1 question · 1 examData Engineering1 question · 1 examSnowflake1 examMicrosoft Fabric1 examMaster Data1 exam

Why it's different from generic coding sites

Algorithm puzzles don't tell you whether someone can build a pipeline.

Real schemas, not toy inputs

Column names and types, a worked example, explicit constraints, and hidden datasets large enough that a naive answer runs out of time.

Graded on more than correctness

Your code runs in a sandbox with no network. The verdict carries runtime, peak memory, shuffle stages and partitions — passing with a collect() looks different from passing properly.

Targeted by company and level

Filter by the loop you are interviewing for and the band you are hired at, then take a timed mock round when you want to know where you actually stand.

Credentials people can check

11 timed exams. Items are sampled from a pool, graded on the server, and a pass issues a credential with a public verification page and a one-click add to your LinkedIn profile.

Data TestDelta LakeLakehouse Engineering on DatabricksMicrosoft Fabric for Data EngineersPySpark DeveloperPython for Data EngineeringSnowflake for Data EngineersSQL for Data EngineeringData Modeling for the LakehouseMaster Data ManagementSpark Optimization

Issued by GeekCoders. Not vendor certifications, and not affiliated with the platforms they cover.

Practice · filtered
Deduplicate customers, keep latest recordPySparkMedium
Find duplicate invoicesSQLEasy
Rolling 7-day active usersSQLMedium
Broadcast join threshold tuningPySparkHard
Skew-safe join on order_idPySparkHard
27 questions · 10 company prep setsBrowse all

Prepare for the level you're hired at

Four tracks, each with its own question pool and timed mock rounds.

0–2 YEARS

Fundamentals

Get past the screening round.

  • Joins and aggregations
  • Python data handling
  • PySpark transformations
2–5 YEARS

Production work

The most common hiring band.

  • Window functions
  • Partitioning and file layout
  • Delta Lake and MERGE
5–8 YEARS

Scale and cost

Where optimization rounds start.

  • Spark tuning and skew
  • Streaming pipelines
  • Lakehouse design
8+ YEARS

Platform and design

Architecture and leadership loops.

  • Distributed systems
  • Governance and cost
  • Multi-cloud platform design

Your next interview is a data engineering interview.

Free forever for core practice. No card required.

Create your accountSee a sample problem
GeekCoders CodeArena

Built by data engineers who got tired of being screened on linked lists.

hello@geekcoders.dev

PracticeLeaderboardCertificationsFeedbackContact

© 2026 GeekCoders CodeArena

Company pages are community-sourced interview preparation material. GeekCoders is not affiliated with, endorsed by, or partnered with the companies or platforms named unless stated.