Mercor · Research profile

Measure what actually matters.

Mohamed A M Elansary, PhD — a decade of multi-model forecast evaluation and uncertainty quantification, now building production agentic evaluation sets, applied to benchmark and dataset design for economically valuable professional work.

Multi-model evaluationUncertainty quantificationScientific data + HPCAgentic evaluation sets

Scientific evaluation

  • Designed multi-model, multi-basin forecast comparisons across hydroclimates.
  • Quantified uncertainty and validated imperfect USGS, NOAA, and NASA observations.
  • Reported failure modes and regime-dependent performance rather than smoothing results.

Production systems

  • Builds GPT, Claude, and Gemini agent workflows at Vertexium.
  • Maintains regression evaluation sets for production agent behavior.
  • Ships retrieval, routing, tenant isolation, and provenance-aware pipelines.

Evaluation approach

Define what correct completion means for a task; build a small evaluation set with explicit provenance; compare honest baselines; examine failures and why they happen; then write a decision-ready account of what the result does and does not support.

Honest fit boundary

I have not authored a published LLM benchmark, built model-as-judge grading pipelines, or worked in RLHF or AI-safety research. My contribution is a decade of multi-model forecast evaluation, uncertainty quantification, and failure-mode analysis, applied to a new subject-matter domain.

Role and location

Research Scientist, APEX Benchmarks · "We work in-person five days a week in our San Francisco, NYC, or London offices." Based in Dallas–Fort Worth and honestly willing to relocate to San Francisco; would discuss the posting's own relocation and housing bonuses as part of that move.

Compensation: Base Salary $200K – $500K • Offers Equity (official posting). Official role posting