Jacob Epifano, Ph.D.

Senior Research Scientist

Machine learning researcher and engineer. Published work on the failure modes of influence-function interpretability and on scaling Bayesian neural networks. Now build production evaluation and observability infrastructure for LLM systems, including the measurement pipelines that distinguish real improvement from apparent improvement, and use them to drive model and prompt iteration. Independent work on singular learning theory, most recently on the conditions under which local learning coefficient estimates do and do not measure what they claim. Looking to apply evaluation and research infrastructure engineering to alignment.

Print this page for a PDF copy.

Research

2026
How the local learning coefficient scales with width in the lazy and rich learning regimes, set by NTK and muP parametrization. A network linear in its parameters does not look regular under the LLC: its Jacobian rank is capped by the dataset rather than the parameter count, so the NTK network's LLC is nearly width-invariant while the muP network's falls with width, ending about 4x lower at width 1024.
2026
Estimating the LLC means sampling from a posterior held near the trained weights by a spring of stiffness gamma, a number conventionally fixed once and rarely reported. At finite gamma the estimator returns a curvature-weighted count of directions, so fixing gamma across models whose curvature differs by 400x measures each of them differently. Proposes fixing the dimensionless ratio kappa instead and validates the resulting bias against models whose LLC is known exactly.

Experience

Senior Research Scientist
Oct 2025 - Present
Sirona Medical
Evaluation and measurement infrastructure
  • Led the migration of the production ASR provider and built provider trialing, evaluation, and management directly into the application. Converted the incumbent single-vendor transcription service into a multiplexer that runs several providers against the same audio and returns shadow results, making head-to-head comparison of latency and WER possible on live production traffic rather than offline samples. Cut word-by-word latency by 50% and reduced WER by 20%.
  • Designed and shipped end-to-end evaluation infrastructure covering every production AI feature: LLM-as-judge scoring, continuous sampling of production traces, and direct dev-to-prod comparison. Validated in production, with judge metrics tracking dev eval scores and no significant drift.
  • Introduced a hallucination-detection metric into the evaluation pipeline and used it as the optimization target for prompt and model iteration; auto-impression hallucination rate fell from 25% to 9%.
  • Led an observability program consolidating evaluation, monitoring, and usage tracking across all AI features, enabling apples-to-apples comparison of foundation models and prompts on fixed datasets, turning model selection into a measured decision rather than a judgment call and accelerating release cadence.
  • Prototyped and benchmarked three ASR accuracy-improvement approaches against the same evaluation harness on proprietary medical datasets (denoising BERT, GPT-style post-correction, classical language-model rescoring).
  • Built Datadog dashboards tracking AI output quality and aggregate feature usage, giving the team its first real-time view of production model behavior.
Systems and agentic infrastructure
  • Cut auto-impression inference latency ~87% at p50 (8.0s to 1.0s) and ~61% at p95 (15.6s to 6.09s) through prompt and model optimization.
  • Shipped cross-stack contributions spanning frontend, LLM inference infrastructure (llm-plex), ASR infrastructure (asr-plex), and agentic services (rhino-api).
  • Stood up Epic EHR sandbox integration with an internal agent, validating APIs and mapping accessible data to expand the agent's tool surface; scoped a second integration against Cerner.
Senior Machine Learning Scientist
Sep 2024 - Oct 2025
Altea Healthcare
  • Led a team of 6 ML engineers across four production generative AI systems, defining architecture and product direction while remaining hands-on.
  • Built a medical information retrieval agent on LangGraph orchestration with LangSmith evaluation and RAG over vector databases; 3.5x improvement in customer satisfaction and 2x usage growth over the prior version.
  • Established end-to-end Azure MLOps: CI/CD, model monitoring, and A/B prompt testing for compliant, repeatable deployment.
  • Built a transformer speech-to-text and LLM structured-generation pipeline producing encounter documentation, reducing encounter duration by 50%.
  • Built a real-time note classifier (embeddings + transformer heads) surfacing critical patient conditions, and an automated RAG ingestion pipeline over patient documents.
Advanced Machine Learning Engineer
May 2023 - Sep 2024
CACI Inc.
  • Collaborated on research in deepfake detection and AI-generated-text detection supporting trust and safety initiatives.
  • Built and productionized variational autoencoder models for anomaly detection in network traffic, improving detection accuracy while reducing false positives.
  • Designed and deployed an LLM-based RAG pipeline for automated sanitization of sensitive documents, cutting manual review time by more than 70%.
Senior Data Scientist
Nov 2021 - May 2023
AiCure
  • Introduced a cross-attention transformer architecture fusing audio, video, and text for clinical endpoint prediction.
  • Built a multi-model feature extraction pipeline on AWS Batch for large-scale audio and video processing, reducing processing time by more than 50%.
  • Designed an automated data quality assessment pipeline that filtered raw clinical media before modeling, reducing downstream model error by 25%.
Student Informatics Assistant II
Nov 2018 - Jun 2021
Children's Hospital of Philadelphia
  • Implemented influence functions in PyTorch for model interpretability and stability analysis; this became Revisiting the Fragility of Influence Functions.
  • Built and validated an ICU early-warning system predicting mortality risk from patient vital signs, and a pipeline for early detection of acute respiratory distress events in children.

Core skills

Evaluation and Experimentation
Eval harness design, LLM-as-judge scoring, A/B prompt testing, production trace sampling, drift detection, shadow evaluation and benchmarking frameworks, observability and monitoring (Datadog)
Research Engineering and Numerical Methods
Experiments built from scratch in PyTorch; Hutchinson trace estimation, conjugate-gradient solves, Lanczos eigensolvers, SGLD sampling with mixing and acceptance diagnostics, ground-truth validation against analytic cases
Learning Theory and Interpretability
Singular learning theory and learning coefficient estimation, influence functions, Bayesian neural networks, variational inference, uncertainty quantification
Agentic Systems and Tooling
LangGraph, ReAct, LLMCompiler, MCP tool integration, multi-step planning, agentic coding workflows
Modeling
Transformers, speech-to-text and ASR pipelines, multimodal fusion, classical ML
Technical Stack
Python, SQL, PyTorch, HuggingFace Transformers, LangChain/LangGraph/LangSmith; Azure ML, AWS (Batch, Lambda, S3); CI/CD for LLM applications; vector databases (Qdrant, FAISS, Chroma, Pinecone)
Leadership
Led and mentored ML teams of up to 6 engineers; defined architecture and research direction across four shipped production systems

Education

Ph.D. Electrical and Computer Engineering
2019 - 2023
Rowan University
B.S. Electrical and Computer Engineering
2015 - 2019
Rowan University

Publications

Interpretability and Learning Theory
Efficient Scaling of Bayesian Neural Networks
IEEE Access, 2024PDF
Revisiting the Fragility of Influence Functions
Neural Networks (Elsevier), 2023PDFarXiv
Clinical Machine Learning
Deployment of a Robust and Explainable Mortality Prediction Model: The COVID-19 Pandemic and Beyond
Smart Health, 2(1), 14, 2026PDF
A Comparison of Feature Selection Techniques for First-day Mortality Prediction in the ICU
IEEE International Symposium on Circuits & Systems (ISCAS), 2023PDF
Towards an Explainable Mortality Prediction Tool
Machine Learning for Signal Processing (MLSP), 2020PDF
Machine Learning Analysis of Digital Clock Drawing Test Performance for Differential Classification of Mild Cognitive Impairment Subtypes Versus Alzheimer's Disease
Journal of the International Neuropsychological Society (JINS), 2020PDF