I work at the raw → refined → curated layer of data — turning messy sources into clean, reliable, analytics-ready datasets with Medallion Architecture on Databricks. Most of my time lives inside Spark, Delta Lake, and Unity Catalog, with a growing focus on layering Mosaic AI (RAG, embeddings, LLM integration) on top of solid data foundations.
Currently an Azure Integration Intern at Protiviti, working hands-on with Azure Databricks in a live enterprise environment — bridging coursework and real infrastructure.
| Project | Stack | Link | |
|---|---|---|---|
| 1 |
Mosaic AI Sales Data Project
End-to-end sales pipeline on Medallion Architecture with logging, embedding generation, and AI processing via Databricks Mosaic AI.
|
DatabricksMosaic AIEmbeddings | repo ↗ |
| 2 |
Sales Medallion Pipeline Optimization
Performance-tuned sales pipeline with broadcast joins, fact/dimension modeling, Autoloader, and incremental processing.
|
PySparkDelta LakeAutoloader | repo ↗ |
| 3 |
Mosaic AI
Databricks notebooks demonstrating GenAI workflows — Foundation Models, Embeddings, Vector Search, RAG, AI Gateway.
|
RAGVector Search | repo ↗ |
| 4 |
Data Ingestion using Lakeflow Declarative Pipelines
Automated ingestion pipeline using Databricks Lakeflow Declarative Pipelines with Medallion Architecture.
|
LakeflowAutomation | repo ↗ |
| 5 |
HR Hierarchy App
Role-based HR hierarchy application built on Databricks.
|
DatabricksPython | repo ↗ |
| 6 |
Lakeview Dashboard
Interactive sales analytics dashboard on Delta Lake and Medallion Architecture using Databricks Lakeview.
|
Delta LakeDashboards | repo ↗ |