Data Engineer at SimanPro · West Haven, Connecticut

Pipelines that carry
data, and models
that read it.

I build the batch and streaming infrastructure that moves data at scale — Spark, Airflow, dbt, Kafka, Snowflake — and co-author applied machine learning research across green computing, industrial prognostics, financial risk and medical imaging.

focus=
0
Years engineering
0
Project repositories
0
Tests, all passing
0
Co-authored papers
0
IELTS band
What I actually do

Two jobs that turn out to be one

I move data for a living, and I co-author machine learning research on the side. They look separate. They are not — most of what goes wrong in a published model is a data problem, and most of what I fix at work is a data problem.

Ingest it

Batch and streaming pipelines over structured and unstructured sources. Spark, Airflow, dbt, Kafka. Late-arriving events, replay after an outage, and producers that change schema without telling anyone.

Model it

Warehouse and lakehouse design — staging, conformed dimensions, slowly changing dimensions, tested marts. Snowflake and Azure Synapse in production; DuckDB when a model should build in CI without credentials.

🛡

Guard it

Data contracts that fail the build, column-level lineage, and PII classification that follows the data downstream instead of going stale three months after somebody wrote it down.

Research implementations

Papers, made runnable

Reference implementations of methods from papers I have co-authored. Each ships an evaluation protocol that does not cheat and tests that fail when it does. None of them quote an accuracy figure — they ship no dataset, so any number would be a re-run on different data or a copy from a paper this code did not produce. Both are ways of being wrong in public.

green-anomaly-detection · measured

The one exception, because energy accounting is the contribution: every detector reports its joules and grams of CO₂e beside its detection score. Produced by scripts/benchmark.py, three seeds, 6000 flows.

DetectorTierAvg precisionJoulesAP / joule
mahalanobischeap0.2310.141.62
z-scorecheap0.1470.072.29
one-class-svmexpensive0.1829.770.019
isolation-forest-32moderate0.1328.540.016
lofexpensive0.29638.500.008
isolation-forestmoderate0.13225.010.005
elliptic-envelopemoderate0.192148.540.001
Mahalanobis reaches 0.231 average precision for 0.14 joules. LOF reaches 0.296 for 38.50 joules — buying those last 0.065 points costs 275× the energy. Whether that trade is worth taking depends on your alert budget; the point is that you can now see the trade at all.
Data engineering

The infrastructure half

The patterns I reach for in production, written as libraries with the reasoning in the docstrings and a test for every claim. All four run their suites in CI on Python 3.10, 3.11 and 3.12 — no cluster, no credentials.

Experience

Where the work happened

JUN 2026 — PRESENT · REMOTE

Data Engineer

SimanPro Inc.

Batch and streaming ingestion for structured and unstructured corporate data. Productionised ETL and ELT on Spark, Airflow and dbt; enterprise warehouse and lake design; governance, lineage and compliance; real-time streaming over distributed message queues.

AUG 2024 — MAY 2026 · CONNECTICUT

Data Engineer

Capital One

Automated aggregation pipelines in Python, SQL and Airflow feeding regulatory compliance and risk reporting. Snowflake and Azure Synapse optimisation down to micro-partitioning. Real-time ingestion and anomaly monitoring on large-scale card and transaction data. Power BI self-service models for senior stakeholders.

NOV 2020 — DEC 2022

Data Engineer

Mindtree

Python automation frameworks over large historical datasets, cutting processing overhead by 40%. Enterprise SQL ETL, views and stored procedures serving cross-functional analytical layers. Data quality profiling and technical documentation.

2023 — 2025 · CONNECTICUT

MS, Computer Science

University of New Haven

Preceded by a BSc in Computer Science & Engineering at BRAC University, Dhaka, with a thesis on voice impersonation detection using LSTM-based RNNs and LIME explanations.

Toolbox

What I work with

Languages

PythonSQLJavaC++JavaScriptDartPHP

Pipelines & streaming

Apache SparkAirflowKafkadbtETL / ELTDocker

Warehouses

SnowflakeAzure SynapseBigQueryPostgreSQLDuckDBMySQL

Cloud

AWS S3EC2AthenaAzureFirebase

Machine learning

PyTorchscikit-learnNumPypandasGrad-CAMLIME

Analytics & web

Power BITableauDjangoReactNode.jsFlutter
Publications

Co-authored research

Listed as they appear on Google Scholar. Where a venue or year could not be verified from a primary source it is omitted rather than guessed.

Get in touch

Let's talk data

Open to conversations about data platform work, streaming architecture, and applied machine learning collaborations.