Contents

[OTS RL Environment] - Agentic Private Equity

Sample & Overview

This proposal outlines Sherpa Labs' long-horizon finance RL environments spanning the breadth of megafund private equity work. Five handcrafted worlds are already built, and 3 sample tasks from the first world are ready with full rubrics, golden responses, and model benchmarks. The proposed program scales this foundation to a projected 750+ tasks across the 5 worlds. Each ultra-realistic world simulates an end-to-end private equity transaction and its associated workflow, which can span six months or more.

Tasks in our worlds are anchored to checkpoints, defining the work required to advance the transaction from one stage to the next, mirroring how analyses and work products compound over the course of a real transaction.

The tasks built into this world require institutional finance knowledge that is not publicly accessible but rather passed down within high finance. This is a boutique set of questions which target problems not covered by existing datasets. Each task, on average, represents 5+ hours of PE associate work plus iterative review by a senior PE investor.

This data supports RLVR, reward modeling, and general RL training with dense rewards and criteria that target intermediary work done by the agent, formatting, and the final answer.

The structure of our worlds supports evaluation at multiple horizons: from bounded, multi-hour PE tasks within a single checkpoint to a single ultra-long-horizon RL episode spanning the full lifecycle of the deal with 100M+ output tokens/600+ PE investment professional hours.

Key Features and Capabilities

Real-World Private Equity: Our data mirrors real PE workflows, which involve complex transaction structuring and messy, heterogeneous data. These workflows rely on institutional knowledge that is not publicly accessible but shared internally within firms and passed down among deal teams. Existing benchmarks and datasets capture textbook vanilla LBOs, which leaves the complexity of real-world deals unsaturated.

Ultra-Long Horizon Worlds: Each world is the culmination of 600+ hours of work exclusively from top-performing associates and VPs at blue chip megafund PE firms, consisting of dozens of artifacts that stretch across dense workbooks, slides, and reports. Each world is split into chronological “scenarios” that represent realistic snapshots of a private equity deal.

Artifacts: Each world is grounded in workbooks of 50+ sheets and hundreds of rows per sheet, interconnected by thousands of formulas. The complexity stems from how artifacts interact: Deal terms such as make-wholes, liquidation preferences, and dividend step-ups can affect calculations throughout the model and trigger cascading calculations across the model. These inputs are scattered across PDFs and spreadsheets in varying formats and granularity. For example, a single assumption pulled from an unstructured term sheet may flow through ownership tracking, into a pro rata tender across 40+ shareholders, and ultimately into return calculations. This requires cross-referencing and reconciliation at every step.

Verifier: Our agentic verifier grades each criterion independently with the same harness and toolkit as the agent model, enabling it to inspect and evaluate complex Excel workbooks. The verifier achieves 96% inter-rater agreement on a criterion level (100% IRA on mechanical criteria, 60% on formatting criteria - n = 87,10) with human annotators who graded the same artifact and rubric by hand.

We separately tested the precision and reproducibility of our grading system through repeated grading runs on a fixed model episode. For each judge we tested, we graded the fixed episode 10 separate times, yielding a total of 970 independent criterion-level grading calls per judge (4,850 in total across judges). Across these repeated evaluations, total scores varied by only 0.9 - 1.2 criteria (one standard deviation) out of 97 criteria, or approximately ±1 criterion per grading. Criterion-level verdict agreement was also consistently high across judges, with pairwise kappa on modal verdicts of 0.79 - 0.94 and pairwise kappa on criterion-level of 0.75 - 0.89.

Together, these results suggest that the verifier is not only well aligned with human judgment but also highly repeatable under a fixed grading regime.

Further details on our verifier validation, including experimental methodology, criterion-level variance, cross-judge agreement, and tests for same-model-family grading bias, are provided in the Appendix.

Harness: We deliberately kept our harness lightweight. It has OpenLibre, PDF, filesystem MCPs, a code exec server, and no web access. Further information is in the Appendix.

Data Spec

Definitions for the private equity environment data specification
TaxonomyDescription
PromptA task that would typically require 5+ hours of a private equity investor’s time, demanding a deep understanding of technical financial concepts and the ability to execute and model them accurately.
prompt_contextSome prompts include a prompt_context field containing additional information required for the task, such as formatting conventions and assumptions. prompt_context is intended to be appended to the end of the prompt.
Data RoomAll artifacts and context that a team of private equity investors can expect at the time of this prompt.
RubricA set of atomic, self-contained, and verifiable conditions that outline all aspects of a golden response; preseeded using a special copilot LLM, then iterated on by our experts with independent rounds of QC.
Golden ResponseA perfect response to the prompt.
VerifierOur agentic LLM grader equipped with the same harness as the agent for evaluating agent-produced outputs.
judge_guidanceIn our experiments, giving the LLM judge a description of the work it is grading provides useful context and leads to far fewer exploratory tool calls. Some criteria are therefore delivered with a judge_guidance field intended for the grader's system prompt. Although it is not strictly required, we encourage its use for more efficient grading.
CategoryHigh-level buckets across core finance principles.
SubcategoryKey financial skills required to execute at elite private equity & financial institutions.
System promptA finance-focused system prompt shared across all tasks.

Delivery: Each task is delivered as a self-contained container image.

Collection Methodology

World Creation and Task Collection: Each world begins with a set of binding constraints such as deal type, company scale, and valuation range, defined by our team alongside senior private equity investors. A first cohort of investors, paired at the associate and vice-president level, builds the source documents that seed the world: financial statements, cap tables, diligence materials, and supporting artifacts, constructed to those constraints and reviewed by a VP before release. A second, independent cohort then works from that source material to build the world forward with the ultimate goal of evaluating the return profile of the transaction; this replicates a deal-team workflow in the real world.

The VP issues a task that requires at least 5+ hours of work to complete, the associate executes, and the two iterate until the output is final, at which point the next task is generated from the state of the deal. At every step, the second cohort produces the golden answer while an agent model attempts the same task in parallel, which is scored using a rubric with 40-120 criteria that carefully describe both quantitative and qualitative aspects of the perfect solution created with the assistance of an LLM copilot.

The result is a world constructed end-to-end by practitioners who do this work daily, carrying their judgment directly to the ultra-realistic environments. Explicit restrictions were enforced around the use of IP-protected data. Artifacts and tasks were run through PII identifiers to ensure no personal or enterprise material non-public data is present.

Chronological Scenarios: Within each world, the transaction is organized into “scenarios”: natural checkpoints in the transaction, each representing a point-in-time containing only the information available at that stage of the deal. These states are then handed over to our team of elite private equity investors who build realistic prompts and golden examples on top of each scenario.

Task Review & Quality Control: Each environment has been through multiple rounds of iteration and audits by a group of VP+ level PE investment professionals. Individual tasks go through an additional 2 rounds of independent review focused on golden response accuracy and rubric comprehensiveness to ensure criteria are reasonable and within expectations.

Category and Subcategory Distribution (Projected)

Transaction Structuring & Valuation

Defines how an acquisition is funded, priced, and structured at purchase.

Builds the sources & uses funding stack, derives implied entry multiples, sets the closing cap table and ownership split, and creates the toggles for the deal scenario that the model evaluates.

Subcategories

  • Sources & Uses / Deal Funding
  • Valuation, Pricing & Transaction Assumptions
  • Pro Forma Capitalization & Ownership
  • Deal Scenarios, Structuring Toggles & Transaction Mechanics

Capitalization & Capital Structure

Tracks how the ownership and debt stack change from entry to exit.

Runs the cap table roll-forward and funds flow, sizes leverage against metric-based grids, schedules amortization, sweep, and interest, and models refinancing and dividend events.

Subcategories

  • Cap Table & Funds Flow Analysis
  • Debt Sizing, Leverage Grids & Debt Schedule
  • Refinancing, Recap & Make-Whole
  • Free Cash Flow, Liquidity & Cash Management

Returns & Exit Analysis

Calculates the investor’s outcome at exit and decomposes the drivers of those returns.

Models full-sale, IPO, and sponsor-recap exit paths to IRR, MOIC, and cash-on-cash, and runs sensitivity analyses on the primary return drivers.

Subcategories

  • Full Sale Returns
  • IPO Sale Returns
  • Sponsor Recap Returns & Dividend Analysis
  • Sensitivity & Returns Grid Analysis

Mergers & Acquisitions Modeling

Handles the financial mechanics of combining or acquiring businesses on top of the existing portfolio company.

Operates bolt-on acquisition engines with target-level financing and pro forma consolidation, quantifies revenue and cost synergies, and computes the incremental ownership and dilution effects of rollover shares, option pools, and earnout instruments.

Subcategories

  • Transformative ComboCo M&A Modeling
  • Bolt-on / Add-on Engines
  • Detailed Revenue & Cost Synergy Modeling
  • Pro Forma Ownership, Dilution & Capitalization Adjustments

5 financial worlds (~100+ artifacts per world) with 150 tasks per world. We have all 5 worlds built and 3 tasks with full rubrics, golden responses, and model benchmarks from the first world. The proposed program extends this to 5 worlds and 750+ tasks, built to buyer specification.

Sample Task

World Setting

[PORTFOLIO COMPANY], a private, PE-backed company, is considering taking a public company, [TARGET COMPANY], private.

Task

Task Context

[PE FIRM] is acquiring [TARGET COMPANY] through [PORTFOLIO COMPANY], its existing portfolio company. Upon closing, [PORTFOLIO COMPANY] and [TARGET COMPANY] combine under ComboCo. The funding comes from: incremental debt, rolled equity, new sponsor cash equity, and balance-sheet cash.

Task prompt

Using the existing model workbook, build the following in the Sources & Uses and Scenario Choose tabs and wire the transaction into the existing model across all four scenarios:

  • Transaction sources & uses, total and cash views (Sources & Uses)
  • Equity value bridges (Sources & Uses)
  • Analysis at various prices for [PORTFOLIO COMPANY], [TARGET COMPANY], and ComboCo (Sources & Uses). Multiples grid for FY2025A, FY2026E, FY2027E.
  • [TARGET COMPANY]: Revenue, Revenue PEG, Adj. EBITDA, Adj. EBITDA (including Run-rate Cost Synergies), Adj. EBITDA (SBC adj.), Unburdened EBITDA, Leverageable EBITDA.
  • [PORTFOLIO COMPANY] multiples: Revenue, Revenue PEG, Adj. EBITDA, Adj. EBITDA (SBC adj.), Unburdened EBITDA, Leverageable EBITDA.
  • ComboCo multiples: Revenue (including Realized Synergies), Revenue PEG (including Realized Synergies), Adj. EBITDA (including Realized Synergies), Adj. EBITDA (including Run-rate Cost Synergies), Adj. EBITDA (SBC adj.) (No Realized Synergies), Adj. EBITDA (SBC adj. including Realized Synergies), Adj. EBITDA (SBC adj. including Run-rate Synergies), Unburdened EBITDA (including Realized Synergies), Leverageable EBITDA (including Realized Synergies).
  • Share calculations and pro forma ownership (Sources & Uses)

Rubric

Exact numerical reference values across all criteria are masked in this public preview to prevent task leakage.

  1. In the delivered state, (which should be transaction scenario offer price ), the sources section should present an Incremental First Lien Term Loan source of (with units of s, with tolerance).
  2. In the delivered state (scenario ), the sources section presents [PORTFOLIO COMPANY] balance-sheet cash used to fund the transaction of (with units of s, with tolerance).
  3. Set the transaction scenario selector to and recalculate: the Incremental First Lien Term Loan source reads (with units of s, with tolerance).
  4. Cycle the transaction scenario selector through and recalculating each time: in every scenario, total sources reads (with units of s, with tolerance) and the Incremental Second Lien source reads .
  5. The live First Lien Term Loan principal is computed by a live formula (not a hardcoded constant) whose input chain includes both the First Lien capacity cap and the residual funding need of the transaction (total uses less cash sources and equity sources).
  6. Set the [TARGET COMPANY] offer price input to (keeping transaction scenario ) and recalculate: the Second Lien Term Loan principal reads exactly (with units of s, with tolerance) - the Second Lien is used only after the First Lien cap is exhausted.
  7. The LTV cap check in the delivered state (scenario ): incremental total debt divided by ComboCo enterprise value at the offer price reads (within absolute), the workbook presents the LTV cap it is tested against, and the associated check/flag section indicates a pass.
  8. Set the transaction scenario selector to and recalculate: [PE FIRM]'s cash proceeds at transaction in the [PORTFOLIO COMPANY] analysis-at-various-prices read (with units of s, with tolerance), which should be half its equity value, since it rolls only .
  9. In the ComboCo multiples grid, Leverageable EBITDA (including Realized Synergies) for FYA at the offer column reads (within ), using a metric of (units of s) minus the ComboCo unburdened cash EBITDA.
  10. The [PORTFOLIO COMPANY] cash source used in the transaction ( in s in the delivered state) is computed by a live formula, not a hard-coded constant, whose input chain reaches the [PORTFOLIO COMPANY] P&L's projected cash balance at the close date (either directly or through the per-scenario assumptions engine).
View the remaining 87 criteria
  1. In the delivered state (scenario ), the Incremental Second Lien Term Loan source reads (s, within absolute).
  2. In the delivered state (scenario ), the sources section presents [TARGET COMPANY] balance-sheet cash used to fund the transaction of (s, within ).
  3. In the delivered state (scenario ), the [PORTFOLIO COMPANY] rollover equity source reads (s, within ).
  4. In the delivered state (scenario ), total sources reads (s, within ). Accept instead only if the workbook presents the convertible settlement gross ( as a use) with the capped-call proceeds of presented as a separate source or offset line.
  5. In the delivered state (scenario ), the uses include a purchase of [TARGET COMPANY] equity of (s, within ), as a single line or a visibly presented subtotal — in scenario nothing is rolled, so the full diluted equity value is purchased for cash.
  6. In the delivered state (scenario ), the uses include the cash settlement of the [TARGET COMPANY] convertible of (s, within ) net of capped-call proceeds; alternatively accept a gross settlement use of with the of capped-call proceeds presented separately as a source or offset.
  7. In the delivered state (scenario ), the uses include the [TUCK-IN ACQUISITION] remaining cash earnout of (s, within ).
  8. In the delivered state (scenario ), the uses include cash to the pro forma balance sheet of (s, within ).
  9. In the delivered state (scenario ), the uses include buyer transaction fees of (s, within ).
  10. In the delivered state (scenario ), the uses include seller transaction fees of (s, within ).
  11. In the delivered state (scenario ), the uses include lender transaction fees of (s, within ).
  12. In the delivered state (scenario ), the cash-only view of the transaction (a separate sources & uses view that excludes rolled equity) shows total cash sources of (s, within ); accept only under the gross-convert presentation (gross settlement of as a use with the capped-call proceeds presented separately).
  13. Set the transaction scenario selector to ('Take Private - Family Rollover') and recalculate: the [TARGET COMPANY] rollover source reads (s, within ).
  14. Set the transaction scenario selector to and recalculate: the use for [TARGET COMPANY] equity purchased for cash (equity value of outstanding, non-rolled shares) reads (s, within ) — the rolled family portion must not be double-counted in the cash purchase.
  15. Set the transaction scenario selector to and recalculate: [PORTFOLIO COMPANY]'s fully diluted pro forma ownership of ComboCo reads (within absolute).
  16. Set the transaction scenario selector to and recalculate: the [TARGET COMPANY] rolling holders' fully diluted pro forma ownership of ComboCo reads (within absolute).
  17. Set the transaction scenario selector to ('Take Private - New GP (Secondary)') and recalculate: the [PORTFOLIO COMPANY] rollover source reads (s, within ) — of [PORTFOLIO COMPANY]'s equity value.
  18. Set the transaction scenario selector to and recalculate: the uses include [PORTFOLIO COMPANY] equity secondary proceeds of (s, within ) — the bought-out half of [PORTFOLIO COMPANY]'s equity is itself presented as a transaction use.
  19. Set the transaction scenario selector to and recalculate: the new sponsor equity source reads (s, within ).
  20. Set the transaction scenario selector to and recalculate: the Incremental First Lien Term Loan source reads (s, within ).
  21. Set the transaction scenario selector to and recalculate: total sources still equals total uses (difference within absolute) and total uses reads (s, within ; accept under the gross-convert presentation with gross use and capped-call proceeds shown separately).
  22. Set the transaction scenario selector to and recalculate: the cash-only view's total cash sources reads (s, within ; accept under the gross-convert presentation) — the new sponsor's secondary purchase money passes to the selling holder and must NOT appear as new cash into the transaction's cash view.
  23. Set the transaction scenario selector to and recalculate: [PORTFOLIO COMPANY]'s fully diluted pro forma ownership of ComboCo reads (within absolute).
  24. Set the transaction scenario selector to and recalculate: the new sponsor's fully diluted pro forma ownership of ComboCo reads (within absolute).
  25. Set the transaction scenario selector to and recalculate: the [TARGET COMPANY] rolling holders' fully diluted pro forma ownership of ComboCo reads (within absolute).
  26. Set the transaction scenario selector to ('Take Private - New GP (Non-Secondary)') and recalculate: the new sponsor equity source reads (s, within ).
  27. Set the transaction scenario selector to and recalculate: the Incremental First Lien Term Loan source reads (s, within ).
  28. Set the transaction scenario selector to and recalculate: [PORTFOLIO COMPANY]'s fully diluted pro forma ownership of ComboCo reads (within absolute).
  29. Set the transaction scenario selector to and recalculate: the new sponsor's fully diluted pro forma ownership of ComboCo reads (within absolute).
  30. Set the transaction scenario selector to and recalculate: the [TARGET COMPANY] rolling holders' fully diluted pro forma ownership of ComboCo reads (within absolute).
  31. Cycle the transaction scenario selector through and recalculating each time: the workbook's global check on the Cover tab (the sum of all page checks) reads (within absolute) in all four scenarios.
  32. Set the transaction scenario selector to and recalculate: the three presented fully diluted pro forma ownership stakes ([PORTFOLIO COMPANY], new sponsor, [TARGET COMPANY] rolling holders) sum to (within absolute).
  33. Set the transaction scenario selector to and recalculate: the cash-only view's total cash sources reads (s, within ; accept only under the gross-convert presentation with the gross settlement as a use and the capped-call proceeds presented separately).
  34. On the per-scenario assumptions engine, the pro forma ComboCo run-rate synergized leverageable EBITDA used for debt sizing reads (s, within ) in the delivered state — this must be the FY figure per the leverage sizing date, not the close-year FY figure of (s).
  35. On the per-scenario assumptions engine, the maximum pro forma total leverage for the live case reads (s, within ) — the sizing EBITDA.
  36. On the per-scenario assumptions engine, the maximum First Lien Term Loan capacity for the live case reads exactly (s, within ) — the sizing EBITDA rounded to the nearest .
  37. On the per-scenario assumptions engine, the maximum Second Lien Term Loan capacity for the live case reads exactly (s, within ) — the residual of maximum total leverage minus First Lien capacity, rounded to the nearest .
  38. Set the [TARGET COMPANY] offer price input to (keeping transaction scenario ) and recalculate: the First Lien Term Loan principal reads exactly (s, within ) — the capacity cap binds at the higher price.
  39. Set the [TARGET COMPANY] offer price input to (keeping transaction scenario ) and recalculate: seller transaction fees read (s, within ) — the investment banking fee is wired to [TARGET COMPANY] enterprise value at the offer price (, rounded to the nearest ) and must move with the price rather than remain at its -price value.
  40. New content follows the layout patterns of comparable sections in the starting file. Sections with no analog, and column choices forced by adjacent content, do not fail.
  41. Outside the region the task asked to fill, everything from the starting file is still where it was. Recalculated values under unchanged formulas do not fail.
  42. Work that the agent does should be formatted the same way across the rest of the workbook.
  43. New sections are bordered the way the workbook borders comparable sections, judged from the rendered page — not from border metadata.
  44. New header bands match the existing layout and structure of the tab's existing ones.
  45. New bordered sections hold their enclosure the way comparable existing sections do, judged from the rendered page.
  46. Toggles and selectors: inputs live where the starting file keeps inputs, and consuming tabs link to them rather than hardcoding.
  47. Long runs of line items are grouped into labeled subtotals.
  48. New check rows look like the workbook's existing checks.
  49. Pre-existing row heights and column widths are unchanged.
  50. In the delivered state (scenario ), the revolving credit facility capacity on the assumptions engine reads (s, within ) — of the First Lien principal.
  51. Set the transaction scenario selector to and recalculate: the revolving credit facility capacity reads (s, within ) — it must track that scenario's smaller First Lien principal.
  52. On the per-scenario assumptions engine, the live-case First Lien Term Loan maturity date reads (exact; years from the close), and the revolver and Second Lien Term Loan maturity dates read the same date.
  53. On the per-scenario assumptions engine, the live-case First Lien closing fee in dollars reads (s, within ) — of the First Lien principal — in the delivered state (scenario ).
  54. On the per-scenario assumptions engine, the live-case First Lien annual amortization in dollars reads (s, within ) — of the First Lien principal — in the delivered state (scenario ).
  55. The [PORTFOLIO COMPANY] equity value bridge presents [PORTFOLIO COMPANY] pre-transaction debt as a deduction of (s, within ; shown negative/as a subtraction, per the reference's sign).
  56. The [PORTFOLIO COMPANY] equity value bridge presents [PORTFOLIO COMPANY] cash as an addition of (s, within ).
  57. In the [PORTFOLIO COMPANY] analysis-at-various-prices, total [PORTFOLIO COMPANY] equity value reads (s, within ) — this AVP bridge additionally deducts sellside fees of ( of TEV) and adds option proceeds of relative to the rollover equity value.
  58. In the [PORTFOLIO COMPANY] analysis-at-various-prices, [PE FIRM]'s fully diluted equity ownership of [PORTFOLIO COMPANY] reads (within absolute).
  59. In the [PORTFOLIO COMPANY] analysis-at-various-prices, [PE FIRM]'s equity value at transaction reads (s, within ).
  60. In the [TARGET COMPANY] analysis-at-various-prices at the offer column, diluted shares outstanding read (s of shares, within ) — basic shares of plus treasury-method dilutive securities plus [TUCK-IN ACQUISITION] stock-earnout shares of .
  61. In the [TARGET COMPANY] dilutive securities calculation, the total treasury-method impact of dilutive securities at the offer price reads (s of shares, within ), built from stock options of at a weighted average strike, RSUs of RSAs of and PSUs of .
  62. the convertible must contribute no shares to the [TARGET COMPANY] diluted share count at any grid price — it is cash-settled at close. Depth: this binds to the line items directly presented in the diluted-share build (options, RSUs, RSAs, PSUs, [TUCK-IN ACQUISITION] earnout shares); a diluted count at of (s, within ) with no convertible share line evidences the exclusion.
  63. In the [TARGET COMPANY] analysis-at-various-prices at the offer column, [TARGET COMPANY] equity value reads (s, within ) — diluted shares times the offer price.
  64. In the [TARGET COMPANY] analysis-at-various-prices at the offer column, [TARGET COMPANY] enterprise value reads (s, within ), bridged from equity value with the convertible debt added back at , the [TUCK-IN ACQUISITION] cash earnout of and cash deducted at (all s).
  65. In the [TARGET COMPANY] analysis-at-various-prices at the current-price column (), the convertible debt add-back reads (s, within ) — the settlement floors at principal at that price (no make-whole, no capped-call value), so the add-back must be price-dependent, not a constant across columns.
  66. In the [TARGET COMPANY] analysis-at-various-prices at the current-price column (), [TARGET COMPANY] enterprise value reads (s, within ).
  67. In the ComboCo analysis-at-various-prices at the offer column, ComboCo enterprise value reads (s, within ) — [TARGET COMPANY] enterprise value at plus the [PORTFOLIO COMPANY] TEV of .
  68. On the assumptions engine / share calculations, the [TARGET COMPANY] family rollover equity value at the offer reads (s, within ) — thousand family shares times sourced from the pre-existing family rollover analysis.
  69. On the per-scenario assumptions engine, the [TARGET COMPANY] rollover percentage of equity value used in the family-rollover scenarios reads (within absolute).
  70. In the [TARGET COMPANY] multiples grid, xRevenue for FYE at the offer column reads (within ), using [TARGET COMPANY] FYE revenue of (s) as the metric.
  71. In the [TARGET COMPANY] multiples grid, xRevenue PEG for FYE at the offer column reads (within ) — the xRevenue multiple divided by times that year's [TARGET COMPANY] revenue growth rate.
  72. In the [TARGET COMPANY] multiples grid, xAdj. EBITDA for FYA at the offer column reads the non-meaningful marker "n.m." (or equivalent text) — the grid must cap multiples above or below as non-meaningful rather than display the raw number.
  73. In the [TARGET COMPANY] multiples grid, xAdj. EBITDA for FYE at the current-price column () reads (within ).
  74. In the [TARGET COMPANY] multiples grid, xAdj. EBITDA (incl. RR Cost Syn.) for FYA at the offer column reads (within ), using a metric of (s) — [TARGET COMPANY] Adj. EBITDA of plus full run-rate cost synergies of .
  75. recalculate the workbook and scan the Sources & Uses and Scenario Choose tabs in each of the four transaction scenario selector states ( ): no cell may contain an error value (#REF!, #VALUE!, #DIV/!, #NAME?, #N/A). Deliberate 'n.a.'/'n.m.' text strings produced by IFERROR or capping logic are not errors.
  76. In the [TARGET COMPANY] multiples grid, the SBC-adjusted Adj. EBITDA metric for FYA reads (s, within ) — a negative figure sourced from the [TARGET COMPANY] P&L spread — and the corresponding xAdj. EBITDA (SBC adj.) multiple at the offer column therefore reads the non-meaningful marker "n.m." (or equivalent text), not a numeric multiple.
  77. In the [TARGET COMPANY] multiples grid, xUnburdened EBITDA for FYA at the offer column reads (within ), using a metric of (s) — [TARGET COMPANY] unburdened (accrual) EBITDA.
  78. In the [TARGET COMPANY] multiples grid, xLeverageable EBITDA for FYA at the offer column reads (within ), using a metric of (s) — the [TARGET COMPANY] unburdened CASH EBITDA, not the accrual unburdened figure of (contamination check between the two EBITDA concepts).
  79. In the [PORTFOLIO COMPANY] multiples grid, xRevenue for FYA reads (within ), using [PORTFOLIO COMPANY] FYA revenue of (s) against the fixed [PORTFOLIO COMPANY] TEV of .
  80. In the [PORTFOLIO COMPANY] multiples grid, xAdj. EBITDA for FYA reads (within ), using a metric of (s).
  81. In the [PORTFOLIO COMPANY] multiples grid, xAdj. EBITDA (SBC adj.) for FYA reads (within ), using a metric of (s) drawn from the [PORTFOLIO COMPANY] P&L's SBC-adjusted EBITDA row.
  82. In the [PORTFOLIO COMPANY] multiples grid, xLeverageable EBITDA for FYA reads (within ), using a metric of (s) — [PORTFOLIO COMPANY] unburdened cash EBITDA.
  83. In the ComboCo multiples grid, xRevenue (incl. Realized Syn.) for FYA at the offer column reads (within ), using ComboCo FYA revenue of (s) against the ComboCo enterprise value at .
  84. In the ComboCo multiples grid, xAdj. EBITDA (incl. RR Cost Syn.) for FYA at the offer column reads (within ), using a metric of (s) — ComboCo Adj. EBITDA plus full run-rate cost synergies of less that year's realized cost synergies of .
  85. In the ComboCo multiples grid, xAdj. EBITDA (SBC adj.) (No Realized Syn.) for FYE at the offer column reads (within ), using a metric of (s) — [TARGET COMPANY] SBC-adjusted EBITDA of plus [PORTFOLIO COMPANY] SBC-adjusted EBITDA of .
  86. In the ComboCo multiples grid, xAdj. EBITDA (SBC adj. incl. Realized Syn.) for FYE at the offer column reads (within ), using a metric of (s) — the no-realized SBC-adjusted metric plus FY realized cost synergies of .
  87. The [PORTFOLIO COMPANY] pre-transaction debt of (s) used in the [PORTFOLIO COMPANY] equity value bridge is computed by a live formula, not a typed constant, whose input chain reaches the [PORTFOLIO COMPANY] P&L's debt balance at the close date (directly or through the per-scenario assumptions engine).

Reference Answer

Explore the workbook that our annotators produced as the golden answer.

Grades

Across 40 independent runs, no model scored above our threshold for a pass. Our experts inspected the result of each episode - no model produced a work product that would be accepted in the real world.

Criterion performance

Model [thinking level]
n = 10
Mean criteria passed / 97Standard deviationMin-maxJudge
Opus 5 [max]0.5380.0790.464-0.701gemini-3.7-flash
GPT 5.6 Sol [ultra]0.5170.0570.423-0.577gemini-3.7-flash
Fable 5 [max]0.4860.0800.351-0.608gemini-3.7-flash
Gemini 3.7 Flash [high]0.3190.0470.237-0.412gpt-5.6-terra
Gemini 3.7 Flash [high]*0.3250.0520.247-0.423sonnet 5

*As a sensitivity check, we regraded the Gemini 3.7 Flash episodes using Sonnet 5 to see if GPT 5.6 Terra was too harsh (Precision - Table 1a). Although the mean score increased by 0.6 pp, Sonnet 5 was unable to rescue Gemini Flash.

Opus 5 records the strongest average result, passing 53.8% of criteria, with GPT 5.6 Sol close behind it at 51.7%. The ranges overlap substantially, however, and the best individual run reaches only 70.1%.

Trajectory & Efficiency Analysis

Run profile

Model [thinking level]
n = 10
Mean turnsMean tool callsMean output tokens
Fable 5 [max]4954173k
Opus 5 [max]127126160k
GPT 5.6 Sol [ultra]2346664k
Gemini 3.7 Flash [high]110109157k

Fable 5 finishes in the fewest turns and tool calls, but in some of the trajectories we inspected, its tool calls pulled in substantially more input tokens per call, reflecting wider searches than those used by Opus 5 or GPT 5.6 Sol. GPT 5.6 Sol takes the most turns but produces the fewest output tokens.

Cost vs. performance

Sources and Uses mean score compared with mean cost per taskGPT 5.6 Sol scores 0.5165 at 8 dollars and 40 cents per task. Opus 5 scores 0.5381 at 13 dollars and 1 cent. Gemini 3.7 Flash scores 0.3186 at 16 dollars and 90 cents. Fable 5 scores 0.4855 at 32 dollars and 47 cents. Vertical error bars show plus or minus one standard deviation in score. Both axes use truncated scales, indicated by axis break marks near their intersection.0.30.40.50.6$8$10$20$40Mean cost per task (USD, log scale)Mean scoreOpus 5: mean score 0.5381, mean cost $13.01 per task, standard deviation 0.079Opus 5$13.01 · 0.538GPT 5.6 Sol: mean score 0.5165, mean cost $8.40 per task, standard deviation 0.057GPT 5.6 Sol$8.40 · 0.516Fable 5: mean score 0.4855, mean cost $32.47 per task, standard deviation 0.080Fable 5$32.47 · 0.485Gemini 3.7 Flash: mean score 0.3186, mean cost $16.90 per task, standard deviation 0.047Gemini 3.7 Flash$16.90 · 0.319
Vertical bars show ±1 standard deviation in score. Scores use 10 runs per model.Pricing note: None of the cost figures reflect promotional pricing.

GPT 5.6 Sol comes closest to the optimal region of this chart: its mean score trails Opus 5 by 2.2 percentage points while costing about 35% less per task. Fable 5 is the most expensive model in this sample despite using the fewest turns and tool calls.

More Sample Tasks

We have more sample tasks to view in more detail - please reach out to justin@sherpalabs.ai for access.

Appendix

Environment Details

Harness: Agents interact with the environment through our standalone, lightweight harness which executes in its own container beside the environment and is only accessible over HTTP. On completion of the task, the environment is sealed for grading.

MCPs: Modified Archipelago servers with custom tools, grouped below.

Code Execution

  • code_exec

OpenLibre Sheets

  • list_tabs_in_spreadsheet
  • read_tab
  • read_csv
  • add_tab
  • delete_tab
  • add_content_text
  • delete_content_cell
  • create_chart
  • filter_tab
  • create_datatable
  • edit_spreadsheet
  • delete_spreadsheet
  • create_spreadsheet

Filesystem

  • list_files
  • read_text_file
  • read_image_file
  • get_file_metadata
  • search_files
  • get_directory_tree

PDF

  • read_image
  • create_pdf
  • read_page_as_image
  • search_pdf
  • read_pdf_pages

Host provided tools

  • recalculate_workbook
  • render_workbook

Compute footprint:

Agent

Harness
2 vCPU / 2 GiB
Environment
4 vCPU / 6 GiB

Grader

Harness
2 vCPU / 1 GiB
Environment
4 vCPU / 6 GiB

Frontier Model Scores

We define a pass as a total score greater than 0.80. Across 10 independent runs per model across our 3 sample tasks - 120 runs overall - no evaluated model produced a passing response. Observed pass@1 was therefore 0% for every model (0/10 per model; 0/120 overall).

We do not report additional pass@k values because, with no observed passes, each point estimate is also zero and provides no additional information.

Our experts noted that most model responses would not even be accepted as a first draft in the real world, which we believe highlights the current unusability of frontier models in high-finance workflows.

Note: With zero observed passes in 10 runs, the two-sided 95% exact confidence interval for each model's pass probability is approximately 0%-30.8%. The observed result should not be read as proof that the underlying probability is exactly zero.

Verifier Testing Methodology and Results

Below, we detail our experiments testing the precision and accuracy of our verifier.

Precision

For each of the 5 judge models, we graded the same agent output (a Fable 5 max-reasoning episode) 10 times against the sample task's 97-criterion rubric to test the precision of our design (for a total of 5 * 10 * 97 = 4,850 individual criteria grades). We report the intra-judge spread below in Table 1a.

To measure inter-rater reliability, we report pairwise kappa at both the modal-verdict and individual-call levels.The call-level measure answers: “If I randomly select one grading call from each judge for the same criterion, how consistently do they agree after adjusting for chance?”

For the modal-verdict kappa, each judge's ten verdicts on a criterion are first collapsed to that judge's majority verdict. Cohen's κ is then computed between each pair of judges over the 97 resulting modal verdicts, with chance agreement estimated from each judge's overall pass rate.

For the call-level kappa, no collapsing occurs. Observed agreement is the mean, across the 97 criteria, of the probability that one randomly drawn grading call from each judge returns the same verdict. It is computed from each judge's per-criterion pass fractions and chance-corrected using marginal pass rates. This is arithmetically equivalent to averaging all 10 x 10 pairings of individual calls per criterion. Because those pairings reuse the same underlying calls, they are not independent observations; uncertainty in κ is assessed by resampling criteria rather than calls.

Table 1a. Per-judge scoring on Single Task Artifact

Judge*Mean criteria passedMean score - k/97 †
Gemini 3.7 Flash46.5648.0% ± 0.9%
Claude Sonnet 544.0045.4% ± 0.9%
Claude Opus 540.5641.8% ± 1.0%
GPT 5.6 Luna39.8041.0% ± 0.8%
GPT 5.6 Terra37.7038.9% ± 0.7%

*Run at high reasoning.

† 95% CI (t-interval).

Table 1b. Pairwise κ over modal verdicts

0.01.0
JudgeGeminiSonnetOpusLunaTerra
Gemini 3.7 Flash0.830.850.870.79
Claude Sonnet 50.830.810.790.83
Claude Opus 50.850.810.890.94
GPT 5.6 Luna0.870.790.890.91
GPT 5.6 Terra0.790.830.940.91

Agreement is consistently high across the judge panel. Opus 5 and GPT 5.6 Terra align most closely, with κ = 0.94.

Table 1c. Pairwise κ over individual grading calls

Each cross-judge estimate averages all call pairings across 97 criteria—100 pairings per criterion. These pairings reuse the underlying grading calls and should not be interpreted as 9,700 independent observations. The diagonal shows each judge's self-agreement across its own repetition pairs, which is the ceiling for that row.

0.01.0
JudgeGeminiSonnetOpusLunaTerra
Gemini0.950.830.830.850.75
Sonnet0.830.910.840.820.82
Opus0.830.840.950.860.89
Luna0.850.820.860.900.87
Terra0.750.820.890.870.92

Cross-judge agreement ranges from 0.75- 0.89, compared with 0.79-0.94 for the modal verdicts in Table 1b.

Note: The diagonal draws without replacement so that each judge's self-kappa is unbiased.

Conclusion

These results lead us to believe that the judges are mostly reading the rubric the same way. All five judges' majority verdicts agree on 84 out of 97 criteria, and all 50 individual grading calls agree outright on 73 out of 97. The remaining disagreements all occur where at least one judge is internally inconsistent, which points to sampling noise rather than a different interpretation of the criterion. We also see no sign of same-family favoritism: every significant judge-by-capsule effect runs in the opposite direction, with same-family judges relatively harsher, and all five judges rank the episode in the same order.

Accuracy

Unlike the repeatability analysis above, we tested whether our verifier is actually right with respect to human graders. We asked 5 of our experts to grade the same Fable 5 max-reasoning episode against an earlier version of the 97-criterion rubric. They agreed unanimously on all 87 criteria checking mechanical correctness, but 2 of the 5 disagreed on 4 of the 10 formatting criteria. We attributed this to those criteria having been written by an expert at a different firm from the two dissenting graders, who both worked at a firm with different formatting conventions. We reworded the criteria to reflect industry-wide standards, added that check to our QC process, and had all 5 experts regrade the same episode against the revised rubric, where they agreed on all 97.

Grading that same revised rubric, Gemini 3.7 Flash agreed with the human consensus on 95.8% of criteria. They disagreed on 3 false positives and 1 false negative - all of which were formatting criteria that relied on exporting the workbook into a PDF and grading visually (and none of which overlapped with the four the experts had originally split on). In other words, for mechanical criteria, human-agentic grader agreement is 100% (87/87), but for formatting criteria, it's only 60% (6/10) - implying that our verifier leans permissive for less-verifiable criteria at this small sample size. We deliver rubrics with mechanical and formatting tags to make criteria-weighting easier.

Formatting is one part of this work that resists verification by construction. Even after ensuring that firm-specific conventions are not encoded in criteria, the industry-wide conventions still have to be judged from a rendered page rather than cell values. Our verifier's errors fall entirely in that bucket. We think verifiable rewards are the right shape for mechanical criteria and preference data is the right shape for the rest, which is why we tag criteria by type and why we can offer our experts' judgement on formatting as RLHF data.

Further Work

Further work will tackle the robustness of our verifier against reward hacking by the agent policy.

Potential Future Data Shapes

We're excited to detail here how we can further extend the richness of our worlds into other data types. Here's just a few things that we have thought about.

Human Preference Training

Many of our experts derided using frontier models in their day-to-day work for unauditable formulas, unusable formatting decisions, and overall poor model structure and integrity. We believe that for models to be widely adopted as copilots or to be fully agentic in financial institutions, the perceived small, fuzzy details matter just as much and need to be done right. We encapsulated some of these into our rubrics, but we can also offer RLHF data or even SFT data.

Advanced Rubric Annotations

  • Weighted criterions - Enable additional signal on critical, nice-to-have, and negative standards. Additional rubric categorizations provide further signal for loss buckets, e.g. formatting, instruction following, formula correctness, value correctness.
  • Graded Responses - Multiple candidate responses per task can be scored using our professional annotators.

Multiturn & Ultra-long Horizon

Since the world is chronological, the work can be split to whatever horizon - tasks can be shorter (e.g. representing a single associate-VP iteration) or a task can be to create the entire multi-month deal. We've also experimented with using another agent to simulate multiturn interactions, such as a “vice president” agent that steers the agent with corrections (much like in real life).

LLM-assisted World Generation

Since our worlds were painstakingly constructed by humans, we believe that it could serve as seed data for synthetic or partially-synthetic environments.