Measure what actually matters.
Mohamed A M Elansary, PhD — a decade of multi-model forecast evaluation and uncertainty quantification, now building production agentic evaluation sets, applied to benchmark and dataset design for economically valuable professional work.
Scientific evaluation
- Designed multi-model, multi-basin forecast comparisons across hydroclimates.
- Quantified uncertainty and validated imperfect USGS, NOAA, and NASA observations.
- Reported failure modes and regime-dependent performance rather than smoothing results.
Production systems
- Builds GPT, Claude, and Gemini agent workflows at Vertexium.
- Maintains regression evaluation sets for production agent behavior.
- Ships retrieval, routing, tenant isolation, and provenance-aware pipelines.
Evaluation approach
Define what correct completion means for a task; build a small evaluation set with explicit provenance; compare honest baselines; examine failures and why they happen; then write a decision-ready account of what the result does and does not support.
Honest fit boundary
I have not authored a published LLM benchmark, built model-as-judge grading pipelines, or worked in RLHF or AI-safety research. My contribution is a decade of multi-model forecast evaluation, uncertainty quantification, and failure-mode analysis, applied to a new subject-matter domain.
Role and location
Research Scientist, APEX Benchmarks · "We work in-person five days a week in our San Francisco, NYC, or London offices." Based in Dallas–Fort Worth and honestly willing to relocate to San Francisco; would discuss the posting's own relocation and housing bonuses as part of that move.
Compensation: Base Salary $200K – $500K • Offers Equity (official posting). Official role posting