You can now delete a finished evaluation run from an agent's Evaluations tab, while a run still in progress cannot be deleted #523
You can now drag the Run column wider on an agent's Evaluations tab to read long run names in full #518
August 2026
Deleting a test no longer warns that it will be removed from other agents, since a test only ever belongs to the one agent it is on #516
Deleting a test from an agent's Tests tab now removes that test from the whole workspace, not just from that agent #512
You can now rename a test run or model comparison so it is easy to find again later #511
You can now click a column heading on a speech-to-text or text-to-speech results table to sort the rows by that column #503
the workspace size limit for a run now applies everywhere a run can be started, including reruns and saved datasets, not only the two buttons that checked it before #390
You can now stop a test run or model comparison while it is still going and keep the results it already collected, and past runs show whether they finished, were stopped, or broke #492
On a simulation run, an evaluator that did not score anything now shows a dash instead of wrongly showing as failed #498
Compare models on a connection agent that has not turned on benchmarking now asks for the model provider and opens the comparison instead of just sending you to the Connection tab, and a failed model row now says "See why" instead of a plain arrow #482
On a shared results link the evaluator columns, scores and the rating chart's scale now show correctly instead of a single generic column or a wrong 0 to 1 scale #496
The evaluations table now resizes its columns to fit what is in them, and every run results page uses the same Results tab name instead of some calling it Outputs #481
On a labelling task's overview, tool call correctness now only shows a card once it actually has a score, so it no longer looks like the other evaluators are missing from the task #494
You can now submit traces where the agent only called a tool and gave no reply for labelling, instead of being told to unpick them first #421
Labelling tasks, evaluators, and bulk upload screens now call this item type "Single agent response" instead of "Agent response" so it matches the agent type it describes #489
Running all tests or comparing models now asks you to confirm first, showing which agent, how many tests, and for a comparison which models will run #484
The top bar of a labelling item no longer changes size when you move between items or switch the live versions only button off #486
A test run now shows a test that could not be run, such as one that timed out or could not reach the judge, as could not be run instead of counting it as a wrong answer #475
The Evaluations tab on an agent no longer shows up until that agent has a run, so you no longer land on a page that just says there is nothing there yet #477
From an evaluator's preview window you can now open its judge prompt for editing in a new tab with one click #474
You can now create a new tool right from the agent's Tools tab, a tool-call test's tool list, or a labelling item's tool-call answer, without leaving to the Tools page first #426
When bulk uploading tests you are now asked what you want to test about the agent, the same question shown when creating a single test #466
people can now filter an agent's traces by label and see that traces can carry labels in the setup code #467
You can now search the annotators list by name, both on the Annotators tab and when assigning annotators to a labelling task #463
On a labelling task, tool call correctness now shows its own card under evaluator scores even before it has a score to show #461
Sending test and benchmark results to a labelling task now carries over the scores the evaluators already gave, so those evaluators do not need to run again on the same text #459
Searching and filtering an agent's tests now covers every linked test instead of just the ones loaded on screen, and removing several tests at once happens in one step #457
Rewrite the changelog lines that were copied from pull request titles #446
A hover box stays open while the pointer is inside it, so the evaluators it lists can be read and opened #445
Find model comparison runs on the agent Evaluations tab with a new filter, and see each run by its name and the models it tried instead of a long id #444
Tick a trace from inside the trace window, and keep your ticks after adding traces to tests or sending them for labelling #443
The buttons on a labelling job page sit beside the score cards again instead of dropping onto a line of their own #433
Links to the Calibrate documentation now open the new documentation site #432
Click an evaluator's name while building a test to see the model, prompt and scores it judges by, without leaving the page #428
Read a long judge prompt in full on the evaluator page, with the View more button back #427
Rows where the agent called a tool are left out of an AI judge run, and every screen now says so instead of showing empty score cards #420
A task on Human alignment shows two evaluators and folds the rest into a plus sign you can hover, so the Evaluators column stays one height #425
A row answered by a tool call is scored by Tool call correctness alone and every other row by the rest, so annotators are no longer asked about evaluators that do not apply #424
Delete is no longer offered on an evaluator that cannot be deleted, such as Tool call correctness #423
See and edit an item's tool calls when adding or editing it, and saving no longer throws them away #422
Send a test where the agent calls a tool for labelling, alongside the tests where it replies with text #417
Send the results of a single agent response test run or model comparison for labelling #416
A single agent response tool call test reads as an input and an output instead of a made-up conversation, and the tool call pass rate sits with the evaluator scores #415
Opening the result of a single agent response test no longer crashes the panel #414
A test's explanation of why it failed shows once, beside the conversation, instead of twice #412
The agents list says whether each agent holds a conversation or gives a single response, single response results read as an input and an output, and the side columns in a results window can be dragged wider #411
The evaluator kind pill is gone from Human alignment, where it named something the reader cannot act on #406
Closing the sidebar sticks across pages and reloads instead of springing open again #410
Creating an evaluator or a labelling task asks what it is for one question at a time, and skips any question with only one possible answer #408
See, create and delete evaluators from an Evaluators tab on the Speech to Text and Text to Speech pages, without starting a run #405
Filter an agent's traces by whether the agent replied or called a tool #403
Say when creating an agent whether it holds a conversation or takes one input and gives one output, and its tests, uploads and evaluators follow that answer #397
Delete a saved version of an evaluator, leaving the current one in place #402
The evaluator page lists its versions down the left and shows one at a time in full, each with edit, delete and Mark as current #400
An evaluator that judges a single response keeps its variables when you add a new version, instead of losing them #401
Every page opened from a list shows the trail of pages leading to it instead of a back arrow #399
The agent Evaluations tab names the evaluators that judged each run #398
A link to a run opens the right window on the page that run is really on, instead of the plain run window on page one #395
The prompt after saving a test says plainly that the evaluator will be added to the agent, and the Evaluators tab updates as soon as you agree #396
The sidebar is shorter, and an agent's past runs move out of the Tests tab into their own Evaluations tab #391
Step between traces with previous and next arrows, or the arrow keys, from inside the trace window #393
The message under a reasoning box sits right under the box instead of below a blank line #389
A labelling job's score row is headed by the name of the person who labelled it, with the job status beside it #387
See what annotators scored beside what the evaluators scored, on the labelling task overview and on a labelling job #386
The assign annotators dialog always keeps its settings in their own column instead of pushing them far below the annotator list #385
Choose when creating a labelling job whether annotators get a comments box, and whether reasoning is optional, required or hidden #384
Upload existing human labels more easily: leave a label blank where nobody scored it, download the items you already have to fill in, and add an annotator without leaving the dialog #382
Read the post announcing Calibrate's talk at IndiaFOSS on the blog #377
A Calibrate link pasted into WhatsApp, LinkedIn or X shows that page's own title, description and picture instead of the home page's #375
Watch the tutorial on connecting your AI tool with Calibrate on the Learn page #373
Read three slide decks on evaluating AI on the Learn page #372
Read why evaluation matters at all in a new Why Calibrate section on the landing page #371
Find every session we have run on evaluating AI, with its recording and its slides, on the new Learn page #370
Read on the landing page how a coding agent can drive Calibrate, with the command to install it #368
Clicking a product area on the landing page puts it in the address, so a reload or a shared link opens the same one #369
Read on the landing page how the conversations your live agent handles turn into tests and human labels #367
The changelog page lists what has changed in the app, and gets a new line every time something ships #365
See how many items an evaluator run has finished while it is still going #364
Duplicate an agent from its own page instead of building a copy by hand #363
Mark an evaluator as optional so annotators can finish an item without scoring it #361
A page you cannot open now says it is blocked instead of showing a number #360
A Google sign-in that does not go through comes back to the Calibrate login page with a message instead of an error page #358
See how many tests each model has finished while a benchmark is running #357
Traces that only made tool calls are left out when you submit traces for labelling, because there is no reply to score #355
Search an agent's traces by anything said in the conversation, the reply, or the trace's own details #354
A trace's extra details now read as a table with room for long values, and the traces list shows its count and page controls above the table #353
Stop the traces setup steps disappearing while you check for traces, which threw away the key you had just created #352
See the conversations your live agent handled on a new Traces tab, open one in full, add them to your tests, or send them for labelling #340
The open test run stays in the address bar on the agent page, so a reload or a shared link opens the same run #347
The item count and page controls sit directly above the items table on a labelling task #349
One click clears a part-filled selection in the labelling dialogs instead of ticking everything first #346
Moving to the next item in a labelling task carries on past the end of the page, and your position counts against every item in the task #348
Evaluators with no human labels yet are marked on the evaluation run page, so a missing agreement number explains itself #343
Send the items your filters leave on screen for review from the labelling job page and the evaluation run page #342
The workspace is now part of the address, so a shared link opens in the workspace it came from #337
Wait on one loading spinner until your workspace is ready instead of seeing the page load in pieces #335
Open your workspace API keys straight from the profile menu or the workspace switcher #334
Sign in from a shared link and land on the page it pointed at instead of the agents list #333
Read Noora Health's own words on the landing page, with a link to their workshop slides #329
Open a link to something kept in another of your workspaces and Calibrate switches to that workspace instead of saying the page is blocked #327
Read the rewritten Kabakoo story on the landing page #325
Keep your item filters after you reload the page or share the link, and see the score cards again on the task overview #320
Filter items by the score an evaluator gave them, on both the evaluation run page and the labelling job page #319
See the task overview even when only evaluators have scored the items and no one has labelled them yet #318
See each evaluator's overall score on the task overview, beside how often it agrees with the annotators #317
See each evaluator's overall score on the evaluation run page, beside how often it agrees with the annotators #312
Read what other teams use Calibrate for on the landing page, with the header staying in view as you scroll #303
The empty gap above the conversation on an evaluation run item is gone #311
Add a new annotator without leaving the assign annotators dialog #310
Each answer option on the annotation card now says what it means #309
Hear the clip before reading the text on a TTS labelling item, with the audio now above the text #302
July 2026
Add custom fields to a connection agent so every request carries them, and give a single test its own values for those fields #291
See on the Connection tab how your agent can report the cost, latency, and tokens of each run so Calibrate can show them #301
An item with every evaluator answered saves itself when the annotator moves to another item, and a part answered one warns before leaving #300
Rank benchmark models by how much you care about cost, quality, and speed, using sliders #298
Annotators now see the whole conversation, including any tool the agent called, when a test run is sent for labelling #295
Compare speech to text and text to speech providers on quality against cost and speed in a new Model selection tab #278
Pick the best value model from a new Top picks tab in benchmark results #286
See what each speech to text and text to speech provider cost for a run, with its own price and currency #277
The Talk to us button now sits at the bottom of the sidebar instead of floating over the page #285
The About tab on speech to text and text to speech results now explains the Sarvam judge scores, and columns with no scores are hidden #282
Check that a connection agent answers before its tests run, with a way to jump to its connection settings when it does not #280
See how long each speech to text provider took, as a new Latency column with its own chart and About entry #274
Stop the tour card and its highlight hanging over the wrong part of the screen while the next step loads #281
Read what each score on a test run or benchmark means in a new About tab #275
Stop a rerun on the tests page opening the old run again #273
Rerun a test run or a benchmark straight from its results window #266
Save a test and run it in one step from the test window #265
The address bar now holds the open test, so a reload or a shared link opens the same test #264
The guided tour now works when the standard evaluators have been renamed or deleted, reusing what is there or creating what it needs #262
New users get a guided tour through their first evaluation #252
See Semantic WER on speech to text results whenever the run works it out, with the judge's reasoning on each row #261
Pick a task on the Human alignment page straight away, instead of waiting while every task is checked first #258
Select several agents on the agents list and delete them in one go #257
Delete speech to text and text to speech evaluations, one at a time or several at once #251
A leaderboard with only one chart keeps it at half width instead of stretching across #250
Run a speech to text evaluation without picking any evaluator #245
Switch on built-in judges that score speech to text transcripts on meaning rather than exact word matches, and read their scores and reasoning in the results #247
Speech to text and text to speech setup is split into Dataset, Models, and Settings tabs so the models are no longer mixed in with everything else #249
The Talk to us button sits at the bottom left of the screen instead of the bottom right #248
See CER next to WER on speech to text results #246
Edit your workspace's default evaluators, add versions to them and delete them, and find them listed under Default instead of among your own #244
Submit rows from a text to speech run for labelling #243
The speech to text and text to speech pickers show only the providers your workspace is set up to use #242
Create a text to speech labelling task, add items with their audio, and have annotators score them #236
Compare the models in a benchmark on a chart of pass rate against cost and speed, and download the chart as an image #241
The view switch above a conversation now sits flush at the top while you scroll through a long conversation #240
Submit speech to text results and simulation run transcripts for labelling, the way test and benchmark results already could be #235
See a spinner while an agent's tabs are still loading, instead of a message saying there is nothing there yet #237
Pick which evaluators an agent uses from a new Evaluators tab on the agent page, and new tests start from those evaluators #231
Changing which evaluators a labelling task uses now saves in one step, so a failure part way through cannot leave the task with the wrong evaluators #234
The Data Extraction tab no longer appears on the agent page #230
Benchmark Gemini for speech to text and text to speech, and see for each provider whether it works in real time or sends the whole audio at once #228
Retry a speech to text or text to speech run that failed before producing any rows, which used to say it could not be retried #229
Evaluator kinds are now called LLM reply and LLM output instead of Conversational reply and LLM response #222
Tests are listed and searched by name only, the description no longer shows under each name #221
Choose how a test search matches the name: contains, starts with, ends with, or exact #219
Filter the tests list by test type: response, tool call, or conversation #220
Groq is no longer offered as a speech to text or text to speech provider #218
Smallest AI text to speech now uses the newer lightning v3.1 voice by default #217
June 2026
Test and benchmark results show a typical response time instead of an average, with the slowest response times noted underneath #213
Large speech datasets upload their audio much faster, every row can still play its clip, and long row numbers no longer overflow their circle #211
Filter by test type when picking which tests to add to an agent #210
Select some tests on the agent Tests tab and compare models on just those tests #209
Test and benchmark results show a separate pass rate for tool call tests instead of folding them into the overall one #208
Switch a test result's conversation between the chat view and its raw text, and copy the raw text in one click #207
Test and benchmark results show which version of the evaluator gave each score #206
Send the results of a test run or benchmark run to a labelling task so annotators can score the same replies #205
Choose which of a task's evaluators annotators are asked to score when you assign them items #203
Create an evaluator that judges a single piece of text rather than a conversation, and build labelling tasks of that kind #201
Accept any value for a tool call setting with the new "Is any" match option, instead of naming the exact value #202
Test and benchmark results show how long each response took and what it cost #200
Choose the evaluators for each row when uploading many tests at once #199
The test count on the agent Tests tab now matches the tests you are actually looking at after filtering, and the test run title has more room #198
A dot next to the run name shows while a test run is still going, the passed and failed counts sit beside it, and long model names in benchmark results are no longer cut short #196
Each test run in the tests list shows how many tests passed, failed, and errored instead of one overall label #195
Search test run and benchmark results by test name, and see tests that hit an error grouped apart from the ones that failed #194
Add several existing tests to an agent in one go, and close the test dialog straight away when you have not changed anything #193
Choose, for each expected tool call argument, whether it must match exactly or be judged by an evaluator #190
Stop the test type picker appearing when you duplicate a test, and stop hover labels sticking on screen after a dialog opens #191
Mark expected tool call arguments as required or optional, and set values that sit inside other arguments or that the tool does not list #189
Switch between the form and a plain text view when adding a tool, so you can paste a whole tool definition at once #188
A page you cannot open now says it was not found or that you do not have access, and switching workspace lands you on the list page for the section you were in #187
Create and revoke keys for a workspace under Workspace settings, so automated runs can reach Calibrate without anyone signing in #184
Stop benchmarks failing to start, or starting twice, when your sign-in was still loading #185
See what each tool returned in tool call test results, next to the tool name and its arguments #183
The landing page now has a Human alignment section showing how human labels are compared with evaluators #181
May 2026
Speech to Text comes first when you pick what kind of labelling task to create #178
The tool call list in the Add test dialog no longer gets cut off at the edge of the dialog #177
Create a Conversation test that checks a whole conversation, and pick the evaluators that score it #172
Duplicate a test or a labelling item straight from its row #173
Select a range of labelling items by holding shift and clicking, and land on Tasks when the Human alignment overview has nothing to show yet #169
Search the items in a labelling task, move through them a page at a time, and keep the page size you chose #168
Drag the evaluators on a labelling task into the order you want annotators to see them #159
Leave a comment on an item while labelling, and read every annotator's comments on that item #160
Download the items in a labelling task as a CSV file, and refresh them without reloading the page #158
Stop the labelling task page showing dashes and jumping to another tab while it is still loading #155
Sort the items in a labelling task by when they were last updated, and the choice is kept for next time #152
See why renaming an annotator failed right under the name box instead of at the top of the page #151
Name the two labels a yes or no evaluator gives, and see those names everywhere its scores appear #148
Move to the next or previous item without leaving the item view, pick all annotators at once, and rename an annotator from the list #136
Narrow an item down to chosen annotators so you can compare just their labels #133
Open any item in a labelling task to see its content, the labels annotators gave it, and what the evaluators scored #131
Switch workspaces from the sidebar, create a new one, and manage who belongs to it #117
Delete labelling jobs one at a time or several at once, and the benchmarking switch on a connected agent saves by itself #116
Add a description to a labelling item, see the time of each turn in a conversation, and upload a file with curly quotes without it failing #107
A connected agent that was never verified keeps your changes as you type, and after you verify it again it asks before saving the new settings #92
Create and edit tests from an agent's Tests tab, filter them by type, run only the ones you pick, and download speech test results as a zip file #88
Share a labelling job or an evaluator run with a link, run an evaluator again, and delete several tests at once #78
Filter an evaluator run to the items where the evaluator and the annotators disagreed, and download the run as a spreadsheet #66
Uploading items in bulk marks the rows that match items you already have, in red where that annotator's existing labels will be replaced #65
See how often an evaluator agreed with the annotators, and what each annotator scored, on the evaluator run page #64
Rating buttons while labelling use the evaluator's own scale instead of always 1 to 5, and annotators can see the criteria values for each item #58