DATA ENGINEER
PYSPARK · DATABRICKS · DELTA LAKE
Open to work
INDIA — IST

Utkarsh Shukla

I build data pipelines that don't trust their inputs. Every record is validated, reconciled and traceable — from raw Bronze to analytics-ready Gold — and bad data is stopped at the gate before anyone downstream sees it.

MEET ME IN YOUR WAREHOUSE
> SELECT * FROM utkarsh;
SCROLL

Showreel

0:36

THIRTY-SIX SECONDS ON HOW I WORK — STREAMING, DATA QUALITY GATES,
AND THE RECONCILED MODEL UNDERNEATH.

UTKARSH — REEL '26
TC 00:00:00:00
RAW IN. INSIGHT OUT.
KAFKA01
BRONZE02
SILVER03
GOLD04
BI05
ANALYTICS-READY
SCHEMATIC
1M+ ORDER-STATUS EVENTS → NEAR-REAL-TIME METRICS EVENT-TIME · WATERMARKS · STATEFUL DEDUP · CHECKPOINTS
BAD DATA STOPS HERE.
PASSED GATE 01 · SCHEMA DRIFT
PASSED GATE 02 · NULL-RATE
BLOCKED GATE 03 · REFERENTIAL INTEGRITY
JENKINS · CRITICAL CHECK FAILED — DOWNSTREAM PROCESSING BLOCKED
THREE SOURCES. ONE TRUTH.
E-COMMERCE APIs
STORE POS FILES
RELATIONAL DBs
RECONCILE
DUPLICATES · MISSING SKUs
SCHEMA · CURRENCY
DIM_PRODUCT SCD2
FACT_SALES
DIM_STORE SCD2
100K → 10M RECORDS BENCHMARKED INCREMENTAL · DELTA LAKE · STAR SCHEMA
0:00 / 0:36 ■ 01 / 03   STREAMING

About

01 / 05

I work where data engineering meets data quality. At Cognizant I build SQL validation suites that reconcile healthcare claims and customer data from source to target, and wire those checks into Jenkins so a broken transformation fails the build instead of a dashboard. Outside of that I build end-to-end pipelines on Databricks — streaming order events through Kafka and Spark Structured Streaming, reconciling retail sales across APIs, POS files and databases, and enforcing data contracts before anything moves downstream. The common thread: pipelines that are idempotent, recoverable, and honest about every record they've processed.

Selected Work

02 / 05

End-to-end builds on Databricks and Delta Lake. Each one is about the parts that usually break in production — late data, duplicates, schema changes, and silent quality drift. Happy to walk through the architecture and trade-offs.

    • Event-driven streaming pipeline ingesting 1M+ asynchronous order-status events through Apache Kafka, producing near-real-time metrics for delivery latency and restaurant preparation times.
    • Event-time processing and watermarking for late-arriving events, with stateful de-duplication for idempotent processing across Bronze, Silver and Gold Delta Lake layers.
    • Checkpoint-based fault tolerance and recovery for streaming workloads; Gold outputs organised into optimised tables for downstream analytical queries.
    • Incremental batch platform reconciling sales records from e-commerce APIs, store POS files and relational databases into a unified sales dataset.
    • Star Schema with SCD Type 2 to preserve historical product and store attributes while handling source-system changes.
    • Automated reconciliation rules catching duplicate transactions, missing SKUs, schema inconsistencies and currency mismatches.
    • Tuned Delta Lake workloads and benchmarked pipeline performance on datasets from 100K to 10M records.
    • Reusable data-quality gateway validating datasets against configurable schema and data contracts before downstream consumption.
    • Automated checks for schema drift, null-rate anomalies, referential integrity violations, duplicate records and unexpected volume changes.
    • Aggregated Data Quality Scorecards and structured validation reports for dataset-level reliability.
    • Embedded in Jenkins CI/CD so ETL tests run on every deployment — and downstream processing is blocked when a critical check fails.

Experience

03 / 05

Cognizant NOW

Programmer Analyst
MAY 2025 — PRESENT
Data Validation · ETL
  • Developed and optimised 50+ SQL validation scenarios — record counts, field-level mappings, transformation rules and source-to-target reconciliation — across ETL pipelines processing healthcare claims and customer data.
  • Validated ETL transformation logic and data mappings across source, transformation and target layers, ensuring business rules were applied correctly end to end.
  • Root-caused production data defects across source datasets, transformation logic and target systems, working with development teams through resolution.
  • REST API validation with Postman, verifying JSON/XML responses and data consistency across downstream interfaces.
  • Integrated automated data-quality checks into Maven and Jenkins CI/CD pipelines, running validation during deployments and cutting reliance on manual ETL verification.
SQL / Python / ETL / Postman / REST APIs / Maven / Jenkins / JIRA

KIET Group of Institutions

B.Tech, Computer Science & IT
2021 — 2025
Ghaziabad · CGPA 8.23
CERTIFICATIONS

Toolkit

04 / 05
LANGUAGES
  • Python
  • SQL
  • Java
  • C++
DATA ENGINEERING
  • Apache Spark / PySpark
  • Databricks
  • Delta Lake
  • Apache Kafka
  • Structured Streaming
  • Medallion Architecture
MODELING & QUALITY
  • Star Schema · SCD Type 2
  • Data Contracts
  • Data Validation
  • Incremental Loading
  • ACID Transactions
DATABASES
  • PostgreSQL
  • MySQL
  • SQL Server
  • JDBC
CLOUD & CI/CD
  • Microsoft Azure
  • AWS
  • Jenkins
  • Maven
  • Git / GitHub
TOOLS & PRACTICE
  • Postman · REST APIs
  • Jupyter · VS Code
  • JIRA · Agile Scrum
  • JSON · XML

Contact

05 / 05
HAVE A ROLE, A PIPELINE PROBLEM, OR JUST WANT TO SAY HI?
Let's talk.