// senior data engineer — google cloud & dbt certified_

Mit Dasondi.

I build cloud-native data platforms that turn raw, messy sources into governed, reliable data products — from append-only bronze to multi-million-row fact models, shipped every sprint.

See the work ↓ Resume.pdf ↗
Currently — Coforge · Client: QXO scroll Est. 2020 — 6+ yrs in data
01 / who

Senior Data Engineer with 6+ years building and owning cloud-native, fault-tolerant data platforms on GCP and AWS.

I ship end-to-end data products every sprint — owning the full lifecycle from source ingestion and append-only bronze (Iceberg / BigLake) through dbt dimensional modeling to a governed publish layer, across DEV / UAT / PROD.

Strengths in root-causing data defects, source-to-target reconciliation, and BigQuery cost tuning — fluent in Python, SQL, and AI-assisted engineering workflows.

Mit Dasondi
dbt run --select fact_delivery
Running with dbt=1.11 · target=prod
1 of 1 OK ...................... [PASS in 42s]
7,142,882 rows · 0 dup keys
0+years engineering data
0%fewer pipeline failures
0+dashboards & reports built
0+industries — logistics · real estate · healthcare · manufacturing
02 / experience

Where the data
flowed

/01 Coforge × QXO Senior Data EngineerGCP Lakehouse · last-mile logistics Apr 2026 — Now
  • Own end-to-end data products shipped every sprint — source extraction → append-only bronze → dbt dimensional modeling → governed publish layer, across DEV / UAT / PROD.
  • Designed, built, and own a 7M-row daily-snapshot last-mile delivery (OTIF) fact model in dbt/BigQuery — the single source of truth for dispatch %, on-time %, and delivery-sales KPIs across the full order history (2020–present).
  • Engineered ingestion-to-bronze with Dataproc/Spark: Mincron (Oracle) → Avro with generated manifests → append-only Iceberg bronze exposed to BigQuery as BigLake external tables.
  • Root-caused and eliminated a ~100K duplicate-key defect — an inclusive time-window join boundary, fixed with a half-open interval + merge→insert_overwrite — removing ~95% of duplicates while preserving every business row.
  • Built repeatable source-to-target reconciliation validating dbt output against authoritative Mincron/Tableau reports to within ~1pp.
  • Cut incremental load cost and runtime via partitioning, clustering, and lookback tuning; daily runs on Airflow/Cloud Composer, every change shipped via GitLab CI/CD.
BigQuerydbtDataprocIcebergBigLakeAirflowGitLab CI/CDAvro
/02 Freedom PI Data EngineerSole DE · entire GCP platform 2024 — 2026
  • Sole data engineer — designed, built, and operated the company's entire GCP data platform end to end: every dataset, pipeline, and streaming ingestion.
  • Reduced pipeline failures and delivery delays by 80% with resilient modular-DAG ELT and retry logic; cut incident response time by 60% via real-time Cloud Monitoring alerts.
  • Developed dbt data marts (CoreLogic, IPScape, KiXy) and 10+ Looker Studio dashboards, replacing Excel workflows with consistent, validated metrics.
  • Architected the full-stack pipeline behind the AI Sales Scorecard; enabled document tracking for 10K+ deals with a GCS metadata service and PDF validation.
  • Automated reporting across 50+ datasets (Apps Script + Gmail extraction pipeline); migrated all credentials to Secret Manager; pioneered Mage-AI for modular transforms.
GCPdbtLooker StudioMage-AICloud MonitoringSecret ManagerApps Script
/03 Meditab Data Science EngineerHealthcare · EHR data Dec 2021 — May 2024
  • Spearheaded data modeling and warehousing for EHR transitions from legacy systems to IMS.
  • Refactored database structures and optimized the event cache, cutting server load by ~26%.
  • Replaced a legacy ETL system with a robust SQL-based solution; automated spreadsheet/JSON ingestion with Python.
  • Designed dynamic, physician-centric reports from structured EMR datasets; optimized AWS S3 with lifecycle policies for cost and versioning.
SQLPythonAWS S3EMR / EHRETL
/04 Beekay Junior Data EngineerEnterprise pipelines · AWS Jun 2020 — Dec 2021
  • Designed SQL-based data pipelines for enterprise clients including Aditya Birla and Ambuja Cement.
  • Cut operational cost by 25% replacing slower Python pipelines; improved data accuracy by 15% with a centralized warehouse.
  • Leveraged AWS S3, Kinesis, Firehose, and Lambda for scalable ingestion and transformation.
AWSKinesisLambdaSQLData Warehouse
03 / stack

Tools of
the trade

01

Cloud Platforms

GCP — BigQuery · Dataproc · Cloud Composer · Cloud Run · Scheduler · Monitoring · Logging
AWS — S3 · Kinesis · Firehose · Lambda

02

Data Engineering

ETL/ELT · dbt · dimensional modeling (fact/dim, SCD2) · medallion / lakehouse · Avro · data migration · web scraping

03

Languages & DBs

Python · SQL · PySpark
BigQuery · PostgreSQL · MySQL · MongoDB

04

Orchestration & Delivery

Airflow · GitLab CI/CD · Mage-AI · Fivetran · Artifact Registry · Soda · Cosmos · FastAPI · OpenAPI

05

AI-Assisted Engineering

Claude / Claude Code · LLM prompt engineering · agentic developer workflows

06

Visualization

Looker Studio · Tableau · Power BI — dashboards people actually open

04 / proof
2021

M.Tech — Computer Engineering

AI-driven research on automated change detection using remote-sensing data for Indian defence (BISAG). Best individual research, Charusat 2021.

2019

B.E. — Computer Engineering

Awarded Best Project of the Year; recognized for startup innovation and academic excellence.

Certified

Google Cloud

Associate Cloud Engineer — the platform I ship on daily.

GCP · ACE
Certified

dbt Labs

dbt Certified Developer — modeling, testing, and shipping analytics code.

dbt Developer