Skills
I'm a senior data scientist with a PhD in particle physics (CERN, statistical analysis of collision data) and production experience building AI systems for financial services. The skills below run from classical statistics and deep learning through to agentic AI engineering, all built on the same rigor I learned analyzing petabyte-scale physics data.
Statistical & Machine Learning Methods
I build and evaluate models end to end, from classical tree ensembles and gradient boosting to deep neural network architectures. Rigorous experimental design, calibration, and interpretability (SHAP, LIME, counterfactual explanations, custom methods) matter to me as much as raw accuracy. For messy, real-world data I reach for probabilistic record linkage and entity resolution (Splink).
Also: scikit-learn, PyTorch, TensorFlow/Keras, Hugging Face Transformers · Random Forests, Gradient Boosting, SVMs, ensemble methods (bagging, boosting, voting) · CNNs, RNNs, LSTMs, Transformers, attention mechanisms, GNNs, GANs, VAEs · feature engineering, hyperparameter optimization (grid/random search, Bayesian optimization) · reinforcement learning (policy gradient methods, PPO, reward shaping, action masking, Gymnasium) · counterfactual explanations, calibration curves, reliability diagrams · pandas, NumPy, SQL, Jupyter/JupyterLab/Colab/Marimo, SciPy, PyMC · matplotlib/seaborn/plotly, Streamlit · hypothesis testing, A/B testing, multivariate analysis, PCA/t-SNE, clustering (k-means, DBSCAN, hierarchical), bootstrap/resampling · Bayesian statistics, experimental design · missing-data imputation (multiple imputation, model-based) · spaCy
Physics & High-Energy Data Analysis
The rigor behind my data science comes from a decade in physics. My PhD was petabyte-scale statistical analysis of CMS collision data at CERN — searching for long-lived particles, published in JHEP — after a master's project applying machine learning to ALICE heavy-ion data. I ran Monte Carlo–based dark matter searches (CRESST) and led the international CMS Level-1 Trigger Menu Group, about ten researchers, through LHC Run-3 preparation.
Also: ROOT (CERN), Geant4, CMSSW, HTCondor, WLCG, grid computing · numerical methods, scientific visualization, high-performance computing (HPC), simulation software · particle detection & instrumentation, FPGA programming, data-acquisition systems · heavy-ion physics, quantum mechanics, relativity, nuclear physics, cosmology
Generative AI & Agentic Engineering
I work fluently across the agentic-AI stack: multi-agent orchestration, agent memory and tool use, Model Context Protocol (MCP) servers and clients, and prompting strategies across decoder-only, MoE, and multimodal architectures. Day to day I work inside agentic coding harnesses (Claude Code, Codex, Cursor), where I've authored custom Agent Skills, hooks, and slash commands. I also fine-tune and quantize open models (LoRA/QLoRA) when a hosted API isn't the right fit.
Also: AI agents / autonomous agents, agentic workflows, tool use / function calling · LangChain, LangGraph · encoder-decoder (T5-style) architectures, small language models (SLMs) · few-shot/zero-shot prompting, chain-of-thought prompting, system prompts, prompt templates · instruction tuning, supervised fine-tuning (SFT), PEFT, quantization · JSON-mode/structured output, content filtering, output validation, uncertainty quantification · vision-language models, image generation (DALL-E, Stable Diffusion, Midjourney) · AI code assistants (GitHub Copilot, Kiro, Amazon Q Developer) · spec-driven development, code generation, AI-assisted code review & debugging
Data Privacy & Governance
Anonymization and governance are where formal rigor pays off directly: k-anonymity/l-diversity/t-closeness and differential privacy for provable guarantees, synthetic data generation and PII detection for safe development pipelines, and GDPR-aligned data handling.
Also: data masking, tokenization, data pseudonymization
Software & Systems Engineering
On the engineering side, I write production Python with the same rigor I'd want in a physics analysis pipeline: typed, tested (pytest, mypy), containerized (Docker/Podman), and observable (Prometheus, Grafana, OpenTelemetry) from day one. I run CI/CD end to end, from pull request to monitored deployment, treating security scanning and drift detection as part of shipping, not an afterthought.
Also: Python, C++, C, Bash · pip/pip-tools/uv/Poetry/conda/venv · mypy, pylint, black · pytest, unittest, doctest, test-driven development, unit/integration/end-to-end testing, performance testing, test automation, mocking/patching, coverage.py, tox · type hints, pydantic · Git, GitHub, GitLab, Bitbucket, git workflows (feature branches, GitFlow) · OOP, design patterns (Factory, Decorator), code review, refactoring, performance optimization, debugging & profiling · Docker, Podman, Docker Compose, Dockerfile optimization · GitHub Actions, GitLab CI, Jenkins · model monitoring, data-drift detection · security scanning · AWS (EC2) · parallel/distributed computing (multi-threading, multi-processing), GPU computing (NVIDIA)
Domain Expertise
I apply this stack to financial back-office operations: AI-powered reconciliation systems, exception management and resolution, and workflow automation under regulatory constraints.
Also: Financial Data Analysis, regulatory compliance
Leadership & Communication
I've led an international, cross-disciplinary team at CERN, and I work directly with stakeholders and customers to ship AI products. I write for both expert and non-expert audiences, from peer-reviewed physics papers to conference talks and technical documentation, and mentor junior team members.
Also: Agile, Scrum, Kanban, backlog management, public speaking, LaTeX, technical specifications, user guides, code documentation, technical training