In my current role, I work on homeowners insurance pricing models and the tooling around model development. Most of my modeling work has focused on production pricing models built on large-scale policy data. I process 10M+ policy records with SQL and PySpark in AWS SageMaker, engineer features from internal and third-party data, and use methods like GBM, LASSO, hierarchical clustering, interaction discovery, and GLM ensemble modeling to improve risk segmentation.
A lot of my recent work has also been about making the modeling workflow more reproducible. I refactored notebook-based workflows into modular Python packages and helped automate training, validation, and experiment tracking with SageMaker Pipelines and MLflow. This has made the work easier to rerun, test, and compare across model iterations.
I also build reusable Python libraries for model development. That includes GitLab CI/CD, automated tests, code quality checks, pre-commit hooks, JFrog package publishing, and Snyk dependency scanning. More recently, I implemented a FastMCP server that exposes custom Python tools so AI agents can retrieve and analyze feature-selection outputs during model development.
On the data engineering side, I build PySpark ETL pipelines with SageMaker Processing Jobs to ingest and transform production insurance records from an enterprise data lake. This part of the work has made me more interested in the full path from raw data to modeling pipeline to production-ready tools.