Work

2025

Data Architecture — PrescriberPoint (via Ballastlane)

PrescriberPoint

Designed and implemented scalable data platforms. Built a Snowflake medallion data-lake to streamline ETL workflows and reduce processing time. Developed end-to-end GenAI and RAG architectures for enhanced data retrieval. Engineered agentic AI workflows with LangGraph, graph embeddings, and vector databases for real-time context, and led ETL/ELT pipelines to structure unstructured data.

  • Azure
  • DBT
  • Snowflake
  • FastAPI
  • Dagster
  • LangGraph
  • RAG
  • GenAI
  • Vector
2025

ClearCast — Captech (via Ballastlane)

Captech

Designed and implemented an Emission Reporting System that forecasts future emissions based on historical data. Developed robust unstructured-data extraction processes from complex Excel files, capturing both data and detailed metadata. Used generative AI (Gemini + LangChain) to standardize and structure metadata files, enabling automated generation of accurate emission reporting forms for frontend applications.

  • AWS Step Functions
  • AWS Lambda
  • Gemini
  • Amazon RDS
  • Amazon S3
  • LangChain
  • Flask
2024

Kroger Throughput Forecast (via BCG X)

Kroger

Designed and implemented a Power BI dashboard to empower regional managers with actionable insights into logistics and warehouse costs, segmented by product and region. Leveraged PySpark for data transformation and orchestrated the pipeline with Databricks, integrating data from Azure Blob Storage. Built an automated pipeline to generate customized performance reports, using LangChain and GenAI to create personalized emails comparing product performance across stores and regions for data-driven decision-making.

  • Databricks
  • Azure
  • PySpark
  • LangChain
  • Power BI
2024

Casas Bahia EVA (via BCG X)

Casas Bahia

Developed and deployed predictive pipelines using PySpark and ML models in Databricks to forecast top-selling items based on product seasonality. Designed KPIs for store managers to monitor store and salesperson performance, and leveraged GenAI to deliver customized insights per salesperson. The solution surfaced real-time updates on the most profitable items and sub-items, driving a 5% increase in total sales and a 14% boost in profits. All datasets were cataloged in Azure Data Catalog; interactive Plotly and Power BI dashboards were integrated into a mobile app, with engagement tracked via GA4 and Google BigQuery.

  • Databricks
  • Azure
  • PySpark
  • Plotly
  • Power BI
  • Azure Data Catalog
  • GA4
  • BigQuery
  • GenAI
2024

Pipeline Optimization — Saint-Gobain (via BCG X)

Saint-Gobain

Optimized and refactored existing Argo data pipelines, achieving a 40% reduction in running time by enhancing PySpark code and parameterizing cluster sizes — resulting in $15K in cost savings. Redesigned pipelines to better align with task dependencies and model scheduling, and streamlined data ingestion into PostgreSQL for seamless stakeholder access.

  • Kubernetes
  • Argo
  • Azure
  • PySpark
  • Plotly
  • PostgreSQL
  • AKS
  • Blob Storage
2024

Teknosa (via BCG X)

Teknosa

Designed and implemented an end-to-end data delivery system using FastAPI to provide KPIs for Teknosa’s online store, enabling seamless API-based consumption of key metrics. The architecture leveraged Azure Blob Storage for ingestion, Polars and Pandas for efficient processing, and Airflow for orchestration, with processed data served from MySQL through the API. The API used GenAI to generate customized HTML messages with tailored product recommendations — significantly enhancing engagement while reducing data-processing costs.

  • FastAPI
  • Polars
  • Pandas
  • Azure
  • MySQL
  • Blob Storage
  • Airflow
  • GenAI
  • LangChain
2023

Intelligent Sales Assistant — BCG X

BCG X

Built a mobile application to help salespersons and managers track sales progress and performance metrics, contributing to both backend development and data engineering. Built the backend with Django and transformed raw GA4 data for actionable insights. Designed an S3 data lake for advanced calculations such as product forecasting and sales predictions, using AWS Lambda for transformations and Step Functions for orchestration. Integrated GenAI to extract information from text-based columns, analyzed processed data in AWS Athena, and built an admin panel in ReactJS to manage feature access and permissions.

  • Python
  • Django
  • GA4
  • AWS S3
  • AWS Lambda
  • Step Functions
  • Athena
  • Pandas
  • ReactJS
2023

Gen AI IVR — Santander (via BCG X)

Santander

Engineered Santander’s first GenAI project in Brazil, revolutionizing the IVR (Interactive Voice Response) system to enhance customer satisfaction. The solution used speech-to-text to transcribe customer conversations for real-time analysis, reducing the steps required to resolve queries. Using LangChain and PySpark Streaming, the system generated embeddings from instructional documents in MongoDB and stored them in a vectorized CosmosDB for efficient retrieval. The project achieved a 22% reduction in IVR handling time and a 20% increase in customer satisfaction.

  • LangChain
  • PySpark Streaming
  • MongoDB
  • CosmosDB
  • GenAI
  • Kafka
2022

Distrito — Data Enrichment & SaaS Solution

Distrito

Expanded and enriched a list of enterprises by capturing additional data from external sources. Developed custom web scrapers to extract data from multiple websites, storing responses as JSON in Amazon S3. Used Airflow to orchestrate the pipeline, triggering processing whenever new files landed in the data lake, and leveraged Pandas and PySpark for transformation and enrichment.

  • Airflow
  • BeautifulSoup
  • Pandas
  • AWS S3
  • PySpark
2022

Distrito — Cloud Infrastructure & Data Warehouse

Distrito

Designed and implemented a cloud-based environment on AWS to support growing data needs. Built a data warehouse (Kimball) using MySQL to store enriched enterprise data, shared through APIs and Looker dashboards as part of the SaaS solution. Later enhanced the architecture with a Delta Lake solution, introducing refinement layers to improve data quality and reliability.

  • Data Lake
  • Looker Studio
  • MySQL
  • Delta Lake
  • AWS
2022

Distrito — Data Integration & Automation

Distrito

With multiple data sources (starting with HubSpot), managing the ELT process became complex. Deployed Airbyte on Kubernetes to streamline ingestion and automate the ELT process. Post-ingestion, used EMR with PySpark for transformation and refinement.

  • Kubernetes
  • EKS
  • Airbyte
  • Airflow
  • AWS
  • PySpark
2021

Distrito — CRUD System Development

Distrito

Developed a CRUD system using Flask to allow users to add, update, or delete enterprise records. Deployed on AWS EC2, ensuring scalability and accessibility for end-users.

  • Flask
  • Jinja2
  • Python
  • MySQL