Building data systems that are reliable by design,
maintainable in practice, and useful to the business.
I approach data infrastructure as a systems problem, not a collection of scripts. A pipeline that runs once is not the same as one that runs correctly every day, and most of the real engineering work lives in that gap: catching failures before they reach a report, every transformation, and keeping each layer of the system separated so one change doesn't break three others downstream.
Certified in Huawei HCIA-Big Data (best trainee across 30+ participants, first attempt) and AWS Cloud Practitioner. Computer Science degree, GPA 3.4/4.0.
Increasingly focused on Data Engineering for AI workflows: building and benchmarking vector store pipelines, embedding systems, and retrieval APIs that run in production,not just in a notebook. The same reliability principles apply here, validated ingestion, measurable latency, and clean separation between indexing and serving. Retrieval quality is measured with an offline evaluation harness, not assumed to be good.
Get In TouchA business question becomes a reliable data product through deliberate stages. Select any stage to see the reasoning, goal, principles, and projects that prove it in practice.
Actively looking for Data Engineering, Big Data, and Cloud Data Engineering roles. If you're building data infrastructure, pipelines, or lakehouses, reach out. I respond within 24 hours.
MoTaha AI
Ask me anything about my work