Building production-grade data infrastructure for Kenyan FinTech.
I design streaming pipelines, dimensional warehouses, and M-Pesa/Daraja-integrated systems the way they should work in production - idempotent, tested, containerised, and monitored.
simulated reference metrics from streaming-pipeline reference architecture
Technical approach
I'm a final-year Data Science student on the Data Engineering track, building toward a career in Kenyan FinTech infrastructure - M-Pesa/Daraja ecosystems, pipeline architecture, and the systems that keep transaction data honest.
Ingestion
Resilient ETL/ELT frameworks treated as software - modular, observable, and schema-aware from the first row landed.
Orchestration
Airflow turns interdependent data tasks into deterministic, automated workflows with retry logic and real alerting.
Dimensional Modeling
Star/snowflake schemas that bridge raw events to business value, optimised for fast analytical queries.
Data CI/CD
Automated testing (Pytest, Great Expectations) and deployment pipelines gating every merge - infrastructure held to engineering rigor.
Medallion architecture, click a layer
Raw Daraja/M-Pesa webhook events move through three tiers before reaching an analytical data mart. Click a node to inspect its schema, data contract, and quality gates.
BRONZE
Raw Daraja webhook ingestion
SILVER
PySpark / dbt cleaning
GOLD
Star-schema analytical marts
Daraja Webhook → Kafka topic → PyFlink/PySpark cleansing → dbt models → Postgres star schema → BI
Code from the stack
Representative excerpts from the streaming-pipeline reference build - the orchestration, transformation, ingestion, and infrastructure layers.
Production data systems
Filter by architecture layer. Every project links to a live GitHub repository.
Grouped by platform layer
Hover a skill for hands-on context.
Ingestion & Streaming
Transformation & Storage
Orchestration & CI/CD
Cloud & Infrastructure
Certifications
From the blog
Notes on pipeline design, streaming architecture, and lessons from production.
Testimonials & technical reviews
"Victor's ability to design scalable data pipelines is exceptional. His Airflow implementations are production-ready and well-documented. A great addition to any data team."
"Demonstrates exceptional understanding of cloud infrastructure, CI/CD pipelines, and distributed systems. Eager to learn and collaborate effectively with the team."
"Proficient in ETL architectures, real-time streaming with Kafka, and warehouse modelling. Shows strong fundamentals and brings fresh perspectives to complex problems."
GitHub activity
Pulled live from the GitHub REST API for @Victor-Kipruto-Rop.
Let's build something reliable
Open to Data Engineering internships and roles. I typically reply within 24 hours.
Schedule a technical chat
Book a 15-minute call to talk pipelines, internships, or collaboration - no form-filling required.