Sean
Fallon
Designing and building production Lakehouse solutions on Azure Databricks for the State of New Jersey — medallion architecture, Unity Catalog governance, and federal compliance reporting. Foundation in enterprise Oracle systems. M.S. Computer Science, NJIT.
Technical Skills
Core competencies- Databricks & Apache Spark
- Delta Lake
- Unity Catalog
- Medallion Architecture
- PySpark
- ELT Pipeline Design
- Data Modeling
- Schema Validation
- Data Quality Frameworks
- Oracle
- PostgreSQL & MySQL
- PL/SQL
- Advanced SQL Optimization
- Databricks SQL
- Python
- FastAPI
- REST API Development
- Shell Scripting
- Azure Databricks
- Azure DevOps
- Git
- Docker
- CI/CD
- Unity Catalog Volumes
- Databricks Jobs
- Delta Time Travel
- Data Lineage
- Pipeline Audit Logging
- Scikit-learn
- TensorFlow & PyTorch
- MLflow
- Feature Pipeline Design
Experience
Professional historyBuilt and deployed a production medallion Lakehouse on Azure Databricks supporting HUD HMIS federal reporting and statewide housing analytics for the Division of Housing and Community Resources.
Developed bronze layer ingestion pipelines with PySpark and Delta Lake across all HUD HMIS tables, including schema validation, null primary key enforcement, record hashing, and dual table pipeline audit logging. Built silver layer transformation notebooks handling deduplication, business rule application, fiscal year labeling, and multi-version EAV assessment pivot logic. Replaced legacy R and Quarto analytical workflows with a Databricks native gold layer.
Configured multi-task Databricks Job orchestration with dependency DAGs and validation gate logic, established Unity Catalog governance standards, and authored architecture decision documents and pipeline standards. Partners with program stakeholders to translate federal HUD data standards into technical architecture decisions.
Presented with the 2025 Innovation and Efficiency Award by the New Jersey State Treasurer for contributions to data modernization and operational efficiency.
Architected migration of legacy Oracle and PL/SQL systems toward service-oriented architectures, exposing database logic as REST APIs via Oracle REST Data Services. Built ELT-style pipelines in SQL, PL/SQL, and Python transforming high-volume financial and operational data into business-ready models, and modernized the legacy SQR reporting platform into SQL and API-driven data products.
Developed centralized SQL models and BI views, built Python workflows for ingestion and data quality auditing, and mentored junior developers through code reviews and engineering best practices.
Coordinated patient services using Electronic Medical Records systems in a regulated healthcare environment, analyzing patient data to match individuals with appropriate services. Built foundational skills in structured data systems, compliance, and workflow analysis that informed a transition into government data engineering.
Freelance WordPress consultant (2016–2018): built data-driven website workflows, automated tasks with Python, and integrated analytics tools to track engagement. Email Marketing and Database Intern at Alfa Art Gallery (2015): managed the email marketing database and supported targeted communication campaigns.
Education
Academic backgroundFocus: Machine Learning, Data Engineering, Algorithms, Deep Learning. Built multi-stage ML pipelines involving preprocessing, feature engineering, model training, and evaluation using TensorFlow, PyTorch, Scikit-learn, and Python.
Focus on communication strategy, analytics, and audience engagement. Member, PRSSA.
Let's Talk
Open to conversations about data architecture, engineering challenges, and collaboration. Based in Oakhurst, NJ, working across the New York City metropolitan area and beyond.