I build data pipelines — ingestion, orchestration, delivery.
Statistician (UNICAMP). Three years building the ingestion and orchestration layer — scrapers, RPA bots, API integrations and scheduled workflows that collect, validate and prepare data for analysis.
Stack, by pipeline layer
The tools I work with at each stage, from extracting data at the source through to delivering it somewhere it can be queried.
Extraction from sources with no usable export path: headless browsers, RPA bots and third-party APIs.
Primarily relational. Schema design and data modeling, plus query optimization for performance on large tables.
Star schemas, incremental load logic, and the statistical layer on top: regression, ANOVA and time series.
Scheduling, retry handling and the integration between steps, so workflows run unattended.
Delivery to the consumer, whether that is a dashboard, an API, a scheduled export or an inbox.
Selected work
code and data on GitHubYahoo Finance → Python ETL → SQLite → a self-contained interactive page. A scheduled GitHub Actions workflow reruns it after each market close and commits the refreshed data.
Twelve public job-portal APIs collected daily into one normalized, classified base in SQLite, served through an interactive dashboard — filter thousands of jobs by area, seniority, company, portal and market.
A five-stage pipeline queries Google Maps, opens every result and returns 14 structured fields per business, a sample of reviews and the tracking pixels each site runs. Async by design; runs on demand.
Experience
Built the lead-capture pipeline end to end: n8n and Make workflows that collect, validate and enrich prospect data from multiple channels, then create and assign CRM records in real time through the HubSpot and RD Station REST APIs. Predictive lead scoring and A/B analysis (regression, ANOVA) run on top of it; Power BI reports conversion, pipeline velocity, cost per lead and campaign ROI off the same data.
Automated the price-collection layer with Python and Selenium, cutting collection time by 80%, and merged SAP transactional extracts into the elasticity model senior leadership priced from. Delivered the real-time price tracking and competitive intelligence dashboards on top.
Stood up the market-intelligence scrapers feeding weekly competitive benchmarks, and rewrote the SQL underneath the job-marketplace KPIs for performance and scalability on large datasets.
Automated critical business processes in Automation Anywhere (Automation 360) and built the pipelines carrying RPA output into SQL databases. Deployed PostgreSQL and Node-RED in Docker containers on Azure, improving pipeline orchestration and reliability.
Time-series forecasting for sales planning and inventory optimization, and data mining for consumption patterns — on the relational databases I built and maintained in SQL.
Get in touch.
Based in Campinas, São Paulo, Brazil.
B.Sc. Statistics
University of Campinas (UNICAMP)