Welcome to

PrakharMittal’s

Portfolio

000

Data Engineering · Banking & Financial Services

Data Solution Analyst at IDFC FIRST Bank

Seven years building ETL and Big Data platforms on Spark, Hadoop and AWS — designing pipelines that move enterprise banking data reliably, at scale, and on time.

Mumbai, Indiaprakharmittal1510@gmail.com

Experience

Transforming data into intelligent business solutions.

With over 7 years of experience in Data Engineering and Big Data, I design scalable cloud architectures, optimize enterprise ETL pipelines, and build high-performance analytics platforms using AWS, Apache Spark, Hadoop, PySpark, SQL, and Airflow. My work has helped global financial institutions improve processing efficiency, reduce infrastructure costs, and unlock the full potential of their data.

  1. IDFC FIRST Bank logo

    IDFC FIRST Bank

    Data Solution Analyst

    • Leads end-to-end development of scalable ETL pipelines on AWS EMR, PySpark and Jupyter, processing high-volume banking datasets with ~30% better processing efficiency.
    • Spearheaded the migration of enterprise workloads from AWS S3/EMR to on-prem HDFS, cutting cloud infrastructure cost by ~25% while tightening governance.
    • Orchestrates dependencies through Apache Airflow DAGs, reducing manual intervention by ~70% and improving scheduling reliability.
    • Optimized PySpark transformations — partitioning, caching, join tuning — for a ~35% cut in job runtime and better cluster utilization.
    • Built interactive validation frameworks in Jupyter that reduced data-validation effort by ~40%.
    • Implemented reconciliation logic, data quality checks and monitoring that lifted data accuracy and consistency by ~85%.
    • Partners with business stakeholders on scalable data solutions, accelerating delivery timelines ~20% for critical reporting.
    • Holds ~99.9% SLA adherence across production data pipelines through proactive monitoring and issue resolution.
  2. ABSA · Barclays Africa Group logo

    ABSA · Barclays Africa Group

    Senior Data Engineer — on-site with LTIMindtree

    • Initiated and led the Data Rail Program, moving legacy Teradata warehouses onto a Hadoop architecture and translating SQL to Spark SQL — 45% faster query processing.
    • Managed 12 Big Data engineers across on-site and offshore teams at >98% delivery compliance.
    • Converted Teradata marts and extracts for 12 African countries into Hadoop marts with Hive, Spark and Teradata utilities — 50% faster ingestion, 30% savings on infrastructure and licensing.
    • Built multi-source business and financial analytics pipelines that doubled ABSA's data-driven revenue potential through better decision-making.
  3. LTIMindtree logo

    LTIMindtree

    Specialist — Data Engineering

    • Supported 100+ enterprise-scale ETL applications across the African continent on Hadoop, Teradata, Informatica, TWS Tivoli, CA WADE and Airflow at 99.9% SLA.
    • Led the Informatica 10.1 → 10.4 Virtual Server migration, tuning mappings, sessions and pushdown queries for a 40% gain in processing speed on high-volume banking data.
    • Designed ETL pipelines with data quality rules, partition pruning and join optimizations, improving accuracy 80–85% across reporting and compliance systems.
    • Deployed Python, Spark, Sqoop and Shell scripts integrating structured and semi-structured data into Hadoop for near-real-time fraud detection and customer insight.
    • Refactored legacy models with partitioning, bucketing, clustered tables and a Spark→Tez engine switch — 35% faster queries, 25% better cluster utilization.
    • Drove the move of on-prem Hadoop data marts to Dell ECS object storage via HDFS–S3A DistCp, tiering, lifecycle policies and JCEKS-encrypted credentials — 45% less on-prem storage overhead, eleven-nines durability with geo-replication.
    • Onboarded 10+ business units with ingestion frameworks and Ranger/Atlas governance, lifting enterprise Hadoop/ECS adoption 50%.

Education & research

  • 2026

    Executive MBA, Business Administration & Management

    Symbiosis International University (SCMHRD), Pune

    • Highest CGPA in the batch
    • Student of the Year

    8.25/10

    CGPA

  • 2019

    B.Tech in Information Technology

    SRM Institute of Science and Technology, Chennai

    7.11/10

    CGPA

  • 2026

    Event Study of the Effect of U.S. 2025–26 Tariffs on the Indian Stock Market

    MBA dissertation — hand-collected dataset, event-study methodology on Indian equity prices.

Resume

The full record, on one page.

PDF · 108 KB · Updated July 2026

Capabilities

The stack, measured.

0+

Years in data engineering

0+

Enterprise ETL applications supported

0

African markets migrated to Hadoop

0%

Production SLA adherence

Skill Domain Focus

  • Big Data & CloudSpark · Hadoop · AWS94
  • Data EngineeringETL · Migration · Tuning95
  • DatabasesTeradata · Oracle · SQL88
  • ProgrammingPython · Scala · Shell86
  • Orchestration & GovernanceAirflow · Ranger · Atlas90

Active Tech Stack Heatmap

95
Spark & PySpark
92
Hadoop EcosystemHDFS · Hive · Tez
88
AWSEMR · S3
86
ETL ToolsInformatica · Sqoop
90
OrchestrationApache Airflow

Interactive Project & Skill Filter

IDFC FIRST Bank

Banking ETL Pipelines

Scalable PySpark & AWS EMR pipelines, reducing latency by 30% for high-volume banking datasets.

PySparkAWS EMRAirflowPython

ABSA · Barclays Africa

Legacy Data Rail Program

Teradata to Hadoop migration across 12 countries, achieving 45% faster query processing.

HadoopTeradataPySparkScala

LTIMindtree

Hybrid Cloud Migration

On-prem Hadoop to Dell ECS (S3) object storage transition, reducing storage overhead by 45%.

HadoopAWS EMRInformaticaPython

Data Pipeline & Quality Viewer

S3 LandingExtractValidateTransformReconcileLoad · HivePublish Mart

Airflow dependency graph — the shape of a nightly banking load: extract fans out to validation and transform, reconciles, then publishes.