Independent Data Science Consultant
High-stakes data science for regulated industries
Philip Sarajlic, MD, PhD, is an independent data science consultant specializing in regulated industries to help organizations apply machine learning and artificial intelligence to complex analytical and strategic challenges. His work spans model development, evaluation, implementation, and AI strategy, with an emphasis on technically rigorous and sustainable solutions.

About
Building and evaluating data systems where errors carry substantial real-world consequences
Experience spanning clinical medicine, biomedical research, finance, quantitative analysis, and production-grade data systems. Across each of these fields, the same principle applies, namely, the developed products need to stand up to scrutiny. Results are ensured to be reproducible, interpretable, and documented clearly enough for production use and verification. A model should remain reliable not just in development, but after it leaves the training environment and is put to use in the real world.
Independent Model Validation
Third-party review of a model, pipeline or analysis. How the validation was designed, where leakage and drift could bite, whether the metrics fit the question, how it behaves across subgroups, whether it is calibrated. What comes out is a written opinion that holds up in front of a board, a regulator or a reviewer.
Clinical and Biomedical Prediction Models
End-to-end model development on patient-level data. The process begins with defining the cohort and the target, then feature engineering across heterogeneous sources, interpretable modeling and calibration. Continues with measuring the result against relevant risk scores already in clinical use and plans for drift measurement.
AI Strategy and Governance
Evaluate whether the organization is ready to adopt AI, where it can create measurable value, and whether the right approach is to build, buy, or integrate existing capabilities. The engagement can then extend to responsible AI principles, privacy-preserving workflows, technical controls, governance structures, and phased deployment planning.
Scientific Evidence and Publication
Study design and statistical analysis plans. Manuscripts, with the revisions and reviewer responses that follow. Grant applications.
Technical Due Diligence and Second Opinion
Independent assessment of AI and health technology claims, for boards, investors and acquirers. Does the evidence support the claim? Does the validation hold? And where is the risk actually sitting?
Generative AI in High-Stakes Settings
Retrieval-augmented systems, evaluation harnesses, prompt design, failure-mode analysis, safety guardrails. Plus a straight answer on whether a language model is the right tool at all for the problem at hand.
Track record
A Record of Measurable Impact
Spearheaded the development of ML platforms
Led the development of end-to-end machine learning platforms for classifier training, fine-tuning, validation, and evaluation, supporting deployment across multiple U.S. states and healthcare systems. The platforms standardized and automated key stages of the model development lifecycle, significantly reducing manual effort and shortening classifier development timelines by orders of magnitude while improving reproducibility, consistency, and scalability.
Research contributing to national recommendations
Research findings cited in national guidelines used to support clinical decision-making, contributing to the evidence considered when shaping recommendations for patient care. Experience conducting research capable of informing questions at a broader clinical and policy level, generating value beyond publications alone.
Completed more than 300 independent reviews
Commissioned by journals, conferences and examination committees to assess someone else’s methods and conclusions. Including, insurance and actuarial models, risk scores, clinical prediction.
Integration of large language models
Designed and integrated local LLM infrastructure capable of running quantized models on local server hardware for use throughout the machine learning development lifecycle. The system combined RAG pipelines, structured context retrieval, and automated evaluation harnesses to turn general-purpose language models into grounded, measurable, and domain-aware technical assistants, ensuring the needs of the end-user were met.
Approach
Practical AI and data science for real-world decision-making
Defensible Methods
Transparent assumptions, reproducible workflows, measurable performance, and documentation that can support technical and business decisions.
Regulated-Industry Judgment
A physician-scientist who has handled healthcare data first-hand: the clinical complexity behind it, the research design around it, and what happens downstream when the output is wrong.
Business-Aligned Execution
An honest reading of whether an initiative is ready to move forward, which risks to deal with first, and what credible results would actually take. The answer then gets translated for whoever needs it, whether that is a technical team, a clinician, an executive or a compliance reviewer.
Start with a consultation
A fixed quote within one business day, and a written summary you keep.















