/enterprise/orgs/{org_id}/model-analysis.
Model Analysis currently shows provider-call counts and hard execution limits. Cost estimates are disabled until model pricing data is available.
Before you start
You need:- Proxy sessions captured through the Enterprise LLM Gateway.
- At least one configured and enabled target model.
- A shared operation title or tag on the sessions you want to compare.
- Organization administrator access.
Create a frozen dataset
A frozen dataset keeps the selected source sessions and evaluation settings fixed. Multiple runs can reuse it, which makes later comparisons reproducible.1
Open Model Analysis
In the Enterprise Dashboard, select Model Analysis, then select New frozen dataset.
2
Describe the dataset
Enter a Dataset name. Purpose is optional and can describe the workload or decision you are evaluating.
3
Select proxy sessions
Choose one tag or one title. Optionally narrow the date range, then select the sessions to freeze. Only matching proxy sessions appear.
4
Choose an evaluation
Select Category agreement or Blind LLM judge, configure it, then select Freeze sessions.
Choose an evaluation method
Category agreement
Use Category agreement when each response contains a classification label, such as an intent, sentiment, routing decision, or risk category. Enter the Category JQ path where the label appears in the captured response. The page reads a sample session, shows candidate leaf paths, and can collect labels across all selected sessions. Review the Allowed labels before freezing the dataset. For example, a label might be stored at:The field selector in a result’s Changes view helps you inspect JSON output. It does not change the dataset’s frozen category field or the aggregate agreement score.
Blind LLM judge
Use Blind LLM judge for open-ended responses where exact labels are not enough. Choose an Independent judge model and enter the Judge rubric. The judge receives the original normalized request and randomized Response A/B outputs. It does not receive model identities, latency, cost, or the generated diff. The source model, target models, and judge model must be distinct. Each successful target response adds a judge call. After a run, the Source ↔ target LLM judge comparison summarizes each target model with source wins, target wins, ties, target win share, and judged coverage. These aggregate counts include only results where the blind judge returned a usable verdict.Prepare and run the replay
1
Select target models
On the frozen dataset, select Prepare full run →. Choose one or more enabled target models.
2
Prepare compatibility preflight
Select Prepare compatibility preflight. ReclaimLLM checks every session-target pair without calling a provider.
3
Review the execution limit
Review compatible cases, candidate calls, judge calls, and Maximum attempts. Incompatible cases remain in the total denominator.
4
Confirm the run
Select Confirm full dataset run to start provider requests.
How requests are replayed
Model Analysis supports captured Chat Completions requests. Before replay, ReclaimLLM:- Keeps supported request fields and drops unsupported fields.
- Replaces the captured model with the selected target model.
- Forces non-streaming output.
- Removes credentials.
- Sends the normalized request through the organization’s governed enterprise proxy.
Review results
The run header separates cases into:- Compatible
- Succeeded
- Incompatible
- Failed
- Unscorable
- Proxy request body shows the normalized session input.
- Structured comparison compares JSON output and lets you specify a field in the Changes tab.
- Side by side displays complete source and target outputs.
- Raw diff displays a unified text diff.
- Attempt history shows replay and judge attempts, errors, and latency.
- Blind-judge results show the preferred response plus strengths and weaknesses.