Utkarsh Shukla
I build data pipelines that don't trust their inputs. Every record is validated, reconciled and traceable — from raw Bronze to analytics-ready Gold — and bad data is stopped at the gate before anyone downstream sees it.
> SELECT * FROM utkarsh;
Showreel
0:36
THIRTY-SIX SECONDS ON HOW I WORK — STREAMING, DATA QUALITY GATES,
AND THE RECONCILED MODEL UNDERNEATH.
SCHEMA · CURRENCY
About
01 / 05I work where data engineering meets data quality. At Cognizant I build SQL validation suites that reconcile healthcare claims and customer data from source to target, and wire those checks into Jenkins so a broken transformation fails the build instead of a dashboard. Outside of that I build end-to-end pipelines on Databricks — streaming order events through Kafka and Spark Structured Streaming, reconciling retail sales across APIs, POS files and databases, and enforcing data contracts before anything moves downstream. The common thread: pipelines that are idempotent, recoverable, and honest about every record they've processed.
Selected Work
02 / 05End-to-end builds on Databricks and Delta Lake. Each one is about the parts that usually break in production — late data, duplicates, schema changes, and silent quality drift. Happy to walk through the architecture and trade-offs.
-
- Event-driven streaming pipeline ingesting 1M+ asynchronous order-status events through Apache Kafka, producing near-real-time metrics for delivery latency and restaurant preparation times.
- Event-time processing and watermarking for late-arriving events, with stateful de-duplication for idempotent processing across Bronze, Silver and Gold Delta Lake layers.
- Checkpoint-based fault tolerance and recovery for streaming workloads; Gold outputs organised into optimised tables for downstream analytical queries.
-
- Incremental batch platform reconciling sales records from e-commerce APIs, store POS files and relational databases into a unified sales dataset.
- Star Schema with SCD Type 2 to preserve historical product and store attributes while handling source-system changes.
- Automated reconciliation rules catching duplicate transactions, missing SKUs, schema inconsistencies and currency mismatches.
- Tuned Delta Lake workloads and benchmarked pipeline performance on datasets from 100K to 10M records.
-
- Reusable data-quality gateway validating datasets against configurable schema and data contracts before downstream consumption.
- Automated checks for schema drift, null-rate anomalies, referential integrity violations, duplicate records and unexpected volume changes.
- Aggregated Data Quality Scorecards and structured validation reports for dataset-level reliability.
- Embedded in Jenkins CI/CD so ETL tests run on every deployment — and downstream processing is blocked when a critical check fails.
Experience
03 / 05Cognizant NOW
- Developed and optimised 50+ SQL validation scenarios — record counts, field-level mappings, transformation rules and source-to-target reconciliation — across ETL pipelines processing healthcare claims and customer data.
- Validated ETL transformation logic and data mappings across source, transformation and target layers, ensuring business rules were applied correctly end to end.
- Root-caused production data defects across source datasets, transformation logic and target systems, working with development teams through resolution.
- REST API validation with Postman, verifying JSON/XML responses and data consistency across downstream interfaces.
- Integrated automated data-quality checks into Maven and Jenkins CI/CD pipelines, running validation during deployments and cutting reliance on manual ETL verification.
KIET Group of Institutions
- Databricks Certified Data Engineer Associate2026
- Microsoft Azure Fundamentals (AZ-900)2026
- Amazon Web Services2026
- Gen AI Fluency — Anthropic Academy2026
Toolkit
04 / 05- Python
- SQL
- Java
- C++
- Apache Spark / PySpark
- Databricks
- Delta Lake
- Apache Kafka
- Structured Streaming
- Medallion Architecture
- Star Schema · SCD Type 2
- Data Contracts
- Data Validation
- Incremental Loading
- ACID Transactions
- PostgreSQL
- MySQL
- SQL Server
- JDBC
- Microsoft Azure
- AWS
- Jenkins
- Maven
- Git / GitHub
- Postman · REST APIs
- Jupyter · VS Code
- JIRA · Agile Scrum
- JSON · XML