Managing Consultant - Data Engineering

Building Scalable
Data Ecosystems
for 8+ Years.

Data engineer with 8+ years building scalable data platforms on AWS and Azure. Architected and led development of a modular data framework powering enterprise analytical and ML pipelines. Owns the full lifecycle from data modeling and lakehouse design through API integration, orchestration, and production deployment.

Experience

A timeline of engineering complex data architectures and leading technical transformations.

March 2021 — PRESENT

Managing Consultant

Systems Limited - Lahore, Pakistan logo Systems Limited - Lahore, Pakistan

  • terminal PartnerLinq - Visibility - Data Engineering Lead
  • terminal Led design and development of a modular Data Framework (Python, Pandas, PySpark, Delta Lake) for ingestion, curation, and analytics — enabling data pipelines without deployment cycles.
  • terminal Owned end-to-end solution architecture across the platform — designed features spanning data ingestion, migration, reconciliation, policy frameworks, alerting, automated reporting, visual analytics, and platform-wide enhancements and optimizations.
  • terminal Data-modeled the metastore, Delta Lake, and SQL Server datastore; optimized for query performance and multi-tenancy.
  • terminal Architected REST and GraphQL APIs for Integration and Semantic services, enabling external systems to submit workloads and consume metadata dynamically.
  • terminal Orchestrated multi-stage pipelines with Databricks Workflows, Apache Airflow, and Azure Data Factory; built a dialect-agnostic SQL query builder.
August 2018 — March 2021

Senior Software Engineer

Northbay Solutions - Lahore, Pakistan logo Northbay Solutions - Lahore, Pakistan

  • terminal Migrated legacy SAS code to PySpark and built reusable dependency packages for AWS Glue ETL jobs; developed a Data Lake framework for large-scale ingestion and processing.
  • terminal Built event-driven ETL triggers using Lambda and CloudWatch on S3 events, and API Gateway endpoints for Kinesis Data Streams with DynamoDB tooling.
  • terminal Designed dynamically parallel Step Functions State Machines and built an automated cloud migration framework with CloudFormation IaC.
  • terminal Developed REST APIs with Flask and a customizable XML parser for document processing.
2017

Intern

Mentor Graphics - Lahore, Pakistan logo Mentor Graphics - Lahore, Pakistan

  • terminal Performed QA and test automation for Nucleus RTOS; wrote test scripts in Python and Bash.

Core Technical Stack

Validated Expertise
code_blocks
Python & Bash
database
SQL & NoSQL
bolt
Apache Spark & Databricks
cloud
AWS & Azure Ecosystems
account_tree
Airflow & Orchestration
dataset
Delta Lake / Data Lakes
api
GraphQL & REST APIs
web
Django & Flask
deployed_code
Infrastructure as Code
stream
Kafka & Big Data Engineering
terminal
Docker & Linux
schema
Data Modeling & Architecture

Featured Projects

Flagship projects that I've created from scratch.

dataset Data Engineering

PartnerLinq - Data Framework

A metadata-driven Spark pipeline framework running on Databricks, managing the full data lifecycle across 20+ independently deployable modules. Built around canonical and semantic model layers, it features a JSON DSL engine that compiles recursive operation trees directly to Spark at runtime — enabling fully metadata-defined transformations without per-tenant code changes. Also supports ML orchestration, policy enforcement, lineage tracking, and materialization.

Lakehouse Data Fabric MLOps
Python Apache Spark Delta Lake Azure Data Lake Storage Azure Key Vault SQL Server Kafka MLflow
api System Integration

PartnerLinq - Visibility - Integration API

A multi-tenant orchestration gateway that routes jobs to Databricks pipelines with full tenant-context isolation — resolving per-client credentials from cloud secrets at runtime. Exposes orchestration surfaces for ingestion, canonical runs, serving, extraction, ML training, cloning, and archival. Includes idempotent job submission, process-flow lifecycle management, canonical validation, Kafka event propagation, and Azure Data Factory integration.

Pipeline Orchestration iPaaS API
Python Django REST Databricks SDK Azure Key Vault Azure Data Factory Kafka SQL Server
grid_view Metadata API

PartnerLinq - Visibility - Metadata API

The platform’s single source of truth for business data models, pipeline configuration, run history, and ML experiments. Exposes both GraphQL (composable, typed queries for orchestration consumers) and REST (CRUD endpoints for legacy compatibility) across a multi-schema store spanning metastore, version, model, annotation, and lineage stores.

Lineage API
Python Django Strawberry GraphQL Django REST Framework SQL Server
table_chart Analytics Query Engine

PartnerLinq - Visibility - Query Builder

A metadata-driven SQL generation service that powers interactive grid UIs and planning applications. Takes a grid model — columns, filters, sorts, run context — and builds parameterized SQL at request time. Supports pivot/group-by, CTE chaining, time-series generation, compare-run views, optimistic row locking, Excel export, and planning modules — all driven by feed metadata from the semantic catalog.

Data Application API
Python Django REST PyPika SQL Server

Certifications

Industry-Recognized Certifications
Microsoft Azure Fundamentals (AZ-900)

Microsoft Azure Fundamentals (AZ-900)

Microsoft

2026
AWS Certified Developer — Associate

AWS Certified Developer — Associate

Amazon Web Services

2018 — 2021

Let's build something scalable.

Currently open to consulting opportunities or senior engineering leadership roles focused on high-scale data infrastructure.