We are former founders and CTOs building the environments and evals frontier labs use to measure and push what models can do.

An eval is only as useful as it is real, so we obsess over realism: environments that mirror actual computer work down to the last detail, with success measured the way it would be in production, not approximated.

Team leadership

We build computer-use environments for frontier labs, contribute to Harbor, and review tasks for Terminal-Bench 3. The app clones and desktop images are our own, and our harness, computer-1, runs the same task on Llama, Claude, Gemini, or GPT, which is how we get the model comparisons. Both founders are co-authors of SWE-Marathon, the long-horizon coding benchmark in review for NeurIPS 2026 and reported on the Grok 4.5, Qwen 3.8, GLM-5.3, Kimi K3, and Hunyuan Hy4 scorecards.

  • Portrait of Christopher Settles

    Christopher Settles

    Co-founder & CEO

    Led evaluation for Uber's generative-AI platform and shipped its first computer-use agents.

    At Refresh, leads world generation and fidelity assurance for the computer-use environments: the cloned apps and desktops agents run in, and whether they behave like the real thing.

    Owns task design and the quality bar: what gets built, how it is graded, and whether it ships.

  • Portrait of Sid Santbakshsing

    Sid Santbakshsing

    Head of Operations

    Founded a company before Refresh, and now leads delivery on every engagement.

    At Refresh, recruits and coordinates the domain experts behind every task, and runs verifier QA and the graded model runs.

    Owns the schedule, the quality gates, and the hand-over on every engagement we have shipped.

    • Delivery lead on the web-clone and desktop task sets
    • Coordinates the human gold-trajectory recordings
  • Portrait of Erik Quintanilla

    Erik Quintanilla

    Co-founder & CTO

    Built the automation behind Amazon's production operations and filed web-scraping patents there, then led platform modernization at Capital One and ML at LynkAI.

    At Refresh, owns verifier design: how every task is graded, from the deterministic checks to the visual judge.

    Owns post-training and the harnesses the agents run through, and sources the artists and technical directors who build the Blender references.

We are hiring. If this sounds like your kind of work,

See open roles