Changelog

September 2026

  • You can now delete a finished evaluation run from an agent's Evaluations tab, while a run still in progress cannot be deleted #523
  • You can now drag the Run column wider on an agent's Evaluations tab to read long run names in full #518

August 2026

  • Deleting a test no longer warns that it will be removed from other agents, since a test only ever belongs to the one agent it is on #516
  • Deleting a test from an agent's Tests tab now removes that test from the whole workspace, not just from that agent #512
  • You can now rename a test run or model comparison so it is easy to find again later #511
  • You can now click a column heading on a speech-to-text or text-to-speech results table to sort the rows by that column #503
  • the workspace size limit for a run now applies everywhere a run can be started, including reruns and saved datasets, not only the two buttons that checked it before #390
  • You can now stop a test run or model comparison while it is still going and keep the results it already collected, and past runs show whether they finished, were stopped, or broke #492
  • On a simulation run, an evaluator that did not score anything now shows a dash instead of wrongly showing as failed #498
  • Compare models on a connection agent that has not turned on benchmarking now asks for the model provider and opens the comparison instead of just sending you to the Connection tab, and a failed model row now says "See why" instead of a plain arrow #482
  • On a shared results link the evaluator columns, scores and the rating chart's scale now show correctly instead of a single generic column or a wrong 0 to 1 scale #496
  • The evaluations table now resizes its columns to fit what is in them, and every run results page uses the same Results tab name instead of some calling it Outputs #481
  • On a labelling task's overview, tool call correctness now only shows a card once it actually has a score, so it no longer looks like the other evaluators are missing from the task #494
  • You can now submit traces where the agent only called a tool and gave no reply for labelling, instead of being told to unpick them first #421
  • Labelling tasks, evaluators, and bulk upload screens now call this item type "Single agent response" instead of "Agent response" so it matches the agent type it describes #489
  • Running all tests or comparing models now asks you to confirm first, showing which agent, how many tests, and for a comparison which models will run #484
  • The top bar of a labelling item no longer changes size when you move between items or switch the live versions only button off #486
  • A test run now shows a test that could not be run, such as one that timed out or could not reach the judge, as could not be run instead of counting it as a wrong answer #475
  • The Evaluations tab on an agent no longer shows up until that agent has a run, so you no longer land on a page that just says there is nothing there yet #477
  • From an evaluator's preview window you can now open its judge prompt for editing in a new tab with one click #474
  • You can now create a new tool right from the agent's Tools tab, a tool-call test's tool list, or a labelling item's tool-call answer, without leaving to the Tools page first #426
  • When bulk uploading tests you are now asked what you want to test about the agent, the same question shown when creating a single test #466
  • people can now filter an agent's traces by label and see that traces can carry labels in the setup code #467
  • You can now search the annotators list by name, both on the Annotators tab and when assigning annotators to a labelling task #463
  • On a labelling task, tool call correctness now shows its own card under evaluator scores even before it has a score to show #461
  • Sending test and benchmark results to a labelling task now carries over the scores the evaluators already gave, so those evaluators do not need to run again on the same text #459
  • Searching and filtering an agent's tests now covers every linked test instead of just the ones loaded on screen, and removing several tests at once happens in one step #457
  • Rewrite the changelog lines that were copied from pull request titles #446
  • A hover box stays open while the pointer is inside it, so the evaluators it lists can be read and opened #445
  • Find model comparison runs on the agent Evaluations tab with a new filter, and see each run by its name and the models it tried instead of a long id #444
  • Tick a trace from inside the trace window, and keep your ticks after adding traces to tests or sending them for labelling #443
  • The buttons on a labelling job page sit beside the score cards again instead of dropping onto a line of their own #433
  • Links to the Calibrate documentation now open the new documentation site #432
  • Click an evaluator's name while building a test to see the model, prompt and scores it judges by, without leaving the page #428
  • Read a long judge prompt in full on the evaluator page, with the View more button back #427
  • Rows where the agent called a tool are left out of an AI judge run, and every screen now says so instead of showing empty score cards #420
  • A task on Human alignment shows two evaluators and folds the rest into a plus sign you can hover, so the Evaluators column stays one height #425
  • A row answered by a tool call is scored by Tool call correctness alone and every other row by the rest, so annotators are no longer asked about evaluators that do not apply #424
  • Delete is no longer offered on an evaluator that cannot be deleted, such as Tool call correctness #423
  • See and edit an item's tool calls when adding or editing it, and saving no longer throws them away #422
  • Send a test where the agent calls a tool for labelling, alongside the tests where it replies with text #417
  • Send the results of a single agent response test run or model comparison for labelling #416
  • A single agent response tool call test reads as an input and an output instead of a made-up conversation, and the tool call pass rate sits with the evaluator scores #415
  • Opening the result of a single agent response test no longer crashes the panel #414
  • A test's explanation of why it failed shows once, beside the conversation, instead of twice #412
  • The agents list says whether each agent holds a conversation or gives a single response, single response results read as an input and an output, and the side columns in a results window can be dragged wider #411
  • The evaluator kind pill is gone from Human alignment, where it named something the reader cannot act on #406
  • Closing the sidebar sticks across pages and reloads instead of springing open again #410
  • Creating an evaluator or a labelling task asks what it is for one question at a time, and skips any question with only one possible answer #408
  • See, create and delete evaluators from an Evaluators tab on the Speech to Text and Text to Speech pages, without starting a run #405
  • Filter an agent's traces by whether the agent replied or called a tool #403
  • Say when creating an agent whether it holds a conversation or takes one input and gives one output, and its tests, uploads and evaluators follow that answer #397
  • Delete a saved version of an evaluator, leaving the current one in place #402
  • The evaluator page lists its versions down the left and shows one at a time in full, each with edit, delete and Mark as current #400
  • An evaluator that judges a single response keeps its variables when you add a new version, instead of losing them #401
  • Every page opened from a list shows the trail of pages leading to it instead of a back arrow #399
  • The agent Evaluations tab names the evaluators that judged each run #398
  • A link to a run opens the right window on the page that run is really on, instead of the plain run window on page one #395
  • The prompt after saving a test says plainly that the evaluator will be added to the agent, and the Evaluators tab updates as soon as you agree #396
  • The sidebar is shorter, and an agent's past runs move out of the Tests tab into their own Evaluations tab #391
  • Step between traces with previous and next arrows, or the arrow keys, from inside the trace window #393
  • The message under a reasoning box sits right under the box instead of below a blank line #389
  • A labelling job's score row is headed by the name of the person who labelled it, with the job status beside it #387
  • See what annotators scored beside what the evaluators scored, on the labelling task overview and on a labelling job #386
  • The assign annotators dialog always keeps its settings in their own column instead of pushing them far below the annotator list #385
  • Choose when creating a labelling job whether annotators get a comments box, and whether reasoning is optional, required or hidden #384
  • Upload existing human labels more easily: leave a label blank where nobody scored it, download the items you already have to fill in, and add an annotator without leaving the dialog #382
  • Read the post announcing Calibrate's talk at IndiaFOSS on the blog #377
  • A Calibrate link pasted into WhatsApp, LinkedIn or X shows that page's own title, description and picture instead of the home page's #375
  • Watch the tutorial on connecting your AI tool with Calibrate on the Learn page #373
  • Read three slide decks on evaluating AI on the Learn page #372
  • Read why evaluation matters at all in a new Why Calibrate section on the landing page #371
  • Find every session we have run on evaluating AI, with its recording and its slides, on the new Learn page #370
  • Read on the landing page how a coding agent can drive Calibrate, with the command to install it #368
  • Clicking a product area on the landing page puts it in the address, so a reload or a shared link opens the same one #369
  • Read on the landing page how the conversations your live agent handles turn into tests and human labels #367
  • The changelog page lists what has changed in the app, and gets a new line every time something ships #365
  • See how many items an evaluator run has finished while it is still going #364
  • Duplicate an agent from its own page instead of building a copy by hand #363
  • Mark an evaluator as optional so annotators can finish an item without scoring it #361
  • A page you cannot open now says it is blocked instead of showing a number #360
  • A Google sign-in that does not go through comes back to the Calibrate login page with a message instead of an error page #358
  • See how many tests each model has finished while a benchmark is running #357
  • Traces that only made tool calls are left out when you submit traces for labelling, because there is no reply to score #355
  • Search an agent's traces by anything said in the conversation, the reply, or the trace's own details #354
  • A trace's extra details now read as a table with room for long values, and the traces list shows its count and page controls above the table #353
  • Stop the traces setup steps disappearing while you check for traces, which threw away the key you had just created #352
  • See the conversations your live agent handled on a new Traces tab, open one in full, add them to your tests, or send them for labelling #340
  • The open test run stays in the address bar on the agent page, so a reload or a shared link opens the same run #347
  • The item count and page controls sit directly above the items table on a labelling task #349
  • One click clears a part-filled selection in the labelling dialogs instead of ticking everything first #346
  • Moving to the next item in a labelling task carries on past the end of the page, and your position counts against every item in the task #348
  • Evaluators with no human labels yet are marked on the evaluation run page, so a missing agreement number explains itself #343
  • Send the items your filters leave on screen for review from the labelling job page and the evaluation run page #342
  • The workspace is now part of the address, so a shared link opens in the workspace it came from #337
  • Wait on one loading spinner until your workspace is ready instead of seeing the page load in pieces #335
  • Open your workspace API keys straight from the profile menu or the workspace switcher #334
  • Sign in from a shared link and land on the page it pointed at instead of the agents list #333
  • Read Noora Health's own words on the landing page, with a link to their workshop slides #329
  • Open a link to something kept in another of your workspaces and Calibrate switches to that workspace instead of saying the page is blocked #327
  • Read the rewritten Kabakoo story on the landing page #325
  • Keep your item filters after you reload the page or share the link, and see the score cards again on the task overview #320
  • Filter items by the score an evaluator gave them, on both the evaluation run page and the labelling job page #319
  • See the task overview even when only evaluators have scored the items and no one has labelled them yet #318
  • See each evaluator's overall score on the task overview, beside how often it agrees with the annotators #317
  • See each evaluator's overall score on the evaluation run page, beside how often it agrees with the annotators #312
  • Read what other teams use Calibrate for on the landing page, with the header staying in view as you scroll #303
  • The empty gap above the conversation on an evaluation run item is gone #311
  • Add a new annotator without leaving the assign annotators dialog #310
  • Each answer option on the annotation card now says what it means #309
  • Hear the clip before reading the text on a TTS labelling item, with the audio now above the text #302

July 2026

  • Add custom fields to a connection agent so every request carries them, and give a single test its own values for those fields #291
  • See on the Connection tab how your agent can report the cost, latency, and tokens of each run so Calibrate can show them #301
  • An item with every evaluator answered saves itself when the annotator moves to another item, and a part answered one warns before leaving #300
  • Rank benchmark models by how much you care about cost, quality, and speed, using sliders #298
  • Annotators now see the whole conversation, including any tool the agent called, when a test run is sent for labelling #295
  • Compare speech to text and text to speech providers on quality against cost and speed in a new Model selection tab #278
  • Pick the best value model from a new Top picks tab in benchmark results #286
  • See what each speech to text and text to speech provider cost for a run, with its own price and currency #277
  • The Talk to us button now sits at the bottom of the sidebar instead of floating over the page #285
  • Save an agent with Cmd+S or Ctrl+S #284
  • The About tab on speech to text and text to speech results now explains the Sarvam judge scores, and columns with no scores are hidden #282
  • Check that a connection agent answers before its tests run, with a way to jump to its connection settings when it does not #280
  • See how long each speech to text provider took, as a new Latency column with its own chart and About entry #274
  • Stop the tour card and its highlight hanging over the wrong part of the screen while the next step loads #281
  • Read what each score on a test run or benchmark means in a new About tab #275
  • Stop a rerun on the tests page opening the old run again #273
  • Rerun a test run or a benchmark straight from its results window #266
  • Save a test and run it in one step from the test window #265
  • The address bar now holds the open test, so a reload or a shared link opens the same test #264
  • The guided tour now works when the standard evaluators have been renamed or deleted, reusing what is there or creating what it needs #262
  • New users get a guided tour through their first evaluation #252
  • See Semantic WER on speech to text results whenever the run works it out, with the judge's reasoning on each row #261
  • Pick a task on the Human alignment page straight away, instead of waiting while every task is checked first #258
  • Select several agents on the agents list and delete them in one go #257
  • Delete speech to text and text to speech evaluations, one at a time or several at once #251
  • A leaderboard with only one chart keeps it at half width instead of stretching across #250
  • Run a speech to text evaluation without picking any evaluator #245
  • Switch on built-in judges that score speech to text transcripts on meaning rather than exact word matches, and read their scores and reasoning in the results #247
  • Speech to text and text to speech setup is split into Dataset, Models, and Settings tabs so the models are no longer mixed in with everything else #249
  • The Talk to us button sits at the bottom left of the screen instead of the bottom right #248
  • See CER next to WER on speech to text results #246
  • Edit your workspace's default evaluators, add versions to them and delete them, and find them listed under Default instead of among your own #244
  • Submit rows from a text to speech run for labelling #243
  • The speech to text and text to speech pickers show only the providers your workspace is set up to use #242
  • Create a text to speech labelling task, add items with their audio, and have annotators score them #236
  • Compare the models in a benchmark on a chart of pass rate against cost and speed, and download the chart as an image #241
  • The view switch above a conversation now sits flush at the top while you scroll through a long conversation #240
  • Submit speech to text results and simulation run transcripts for labelling, the way test and benchmark results already could be #235
  • See a spinner while an agent's tabs are still loading, instead of a message saying there is nothing there yet #237
  • Pick which evaluators an agent uses from a new Evaluators tab on the agent page, and new tests start from those evaluators #231
  • Changing which evaluators a labelling task uses now saves in one step, so a failure part way through cannot leave the task with the wrong evaluators #234
  • The Data Extraction tab no longer appears on the agent page #230
  • Benchmark Gemini for speech to text and text to speech, and see for each provider whether it works in real time or sends the whole audio at once #228
  • Retry a speech to text or text to speech run that failed before producing any rows, which used to say it could not be retried #229
  • Evaluator kinds are now called LLM reply and LLM output instead of Conversational reply and LLM response #222
  • Tests are listed and searched by name only, the description no longer shows under each name #221
  • Choose how a test search matches the name: contains, starts with, ends with, or exact #219
  • Filter the tests list by test type: response, tool call, or conversation #220
  • Groq is no longer offered as a speech to text or text to speech provider #218
  • Smallest AI text to speech now uses the newer lightning v3.1 voice by default #217

June 2026

  • Test and benchmark results show a typical response time instead of an average, with the slowest response times noted underneath #213
  • Large speech datasets upload their audio much faster, every row can still play its clip, and long row numbers no longer overflow their circle #211
  • Filter by test type when picking which tests to add to an agent #210
  • Select some tests on the agent Tests tab and compare models on just those tests #209
  • Test and benchmark results show a separate pass rate for tool call tests instead of folding them into the overall one #208
  • Switch a test result's conversation between the chat view and its raw text, and copy the raw text in one click #207
  • Test and benchmark results show which version of the evaluator gave each score #206
  • Send the results of a test run or benchmark run to a labelling task so annotators can score the same replies #205
  • Choose which of a task's evaluators annotators are asked to score when you assign them items #203
  • Create an evaluator that judges a single piece of text rather than a conversation, and build labelling tasks of that kind #201
  • Accept any value for a tool call setting with the new "Is any" match option, instead of naming the exact value #202
  • Test and benchmark results show how long each response took and what it cost #200
  • Choose the evaluators for each row when uploading many tests at once #199
  • The test count on the agent Tests tab now matches the tests you are actually looking at after filtering, and the test run title has more room #198
  • A dot next to the run name shows while a test run is still going, the passed and failed counts sit beside it, and long model names in benchmark results are no longer cut short #196
  • Each test run in the tests list shows how many tests passed, failed, and errored instead of one overall label #195
  • Search test run and benchmark results by test name, and see tests that hit an error grouped apart from the ones that failed #194
  • Add several existing tests to an agent in one go, and close the test dialog straight away when you have not changed anything #193
  • Choose, for each expected tool call argument, whether it must match exactly or be judged by an evaluator #190
  • Stop the test type picker appearing when you duplicate a test, and stop hover labels sticking on screen after a dialog opens #191
  • Mark expected tool call arguments as required or optional, and set values that sit inside other arguments or that the tool does not list #189
  • Switch between the form and a plain text view when adding a tool, so you can paste a whole tool definition at once #188
  • A page you cannot open now says it was not found or that you do not have access, and switching workspace lands you on the list page for the section you were in #187
  • Create and revoke keys for a workspace under Workspace settings, so automated runs can reach Calibrate without anyone signing in #184
  • Stop benchmarks failing to start, or starting twice, when your sign-in was still loading #185
  • See what each tool returned in tool call test results, next to the tool name and its arguments #183
  • The landing page now has a Human alignment section showing how human labels are compared with evaluators #181

May 2026

  • Speech to Text comes first when you pick what kind of labelling task to create #178
  • The tool call list in the Add test dialog no longer gets cut off at the edge of the dialog #177
  • Create a Conversation test that checks a whole conversation, and pick the evaluators that score it #172
  • Duplicate a test or a labelling item straight from its row #173
  • Select a range of labelling items by holding shift and clicking, and land on Tasks when the Human alignment overview has nothing to show yet #169
  • Search the items in a labelling task, move through them a page at a time, and keep the page size you chose #168
  • Drag the evaluators on a labelling task into the order you want annotators to see them #159
  • Leave a comment on an item while labelling, and read every annotator's comments on that item #160
  • Download the items in a labelling task as a CSV file, and refresh them without reloading the page #158
  • Stop the labelling task page showing dashes and jumping to another tab while it is still loading #155
  • Sort the items in a labelling task by when they were last updated, and the choice is kept for next time #152
  • See why renaming an annotator failed right under the name box instead of at the top of the page #151
  • Name the two labels a yes or no evaluator gives, and see those names everywhere its scores appear #148
  • Move to the next or previous item without leaving the item view, pick all annotators at once, and rename an annotator from the list #136
  • Narrow an item down to chosen annotators so you can compare just their labels #133
  • Open any item in a labelling task to see its content, the labels annotators gave it, and what the evaluators scored #131
  • Switch workspaces from the sidebar, create a new one, and manage who belongs to it #117
  • Delete labelling jobs one at a time or several at once, and the benchmarking switch on a connected agent saves by itself #116
  • Add a description to a labelling item, see the time of each turn in a conversation, and upload a file with curly quotes without it failing #107
  • A connected agent that was never verified keeps your changes as you type, and after you verify it again it asks before saving the new settings #92
  • Create and edit tests from an agent's Tests tab, filter them by type, run only the ones you pick, and download speech test results as a zip file #88
  • Share a labelling job or an evaluator run with a link, run an evaluator again, and delete several tests at once #78
  • Filter an evaluator run to the items where the evaluator and the annotators disagreed, and download the run as a spreadsheet #66
  • Uploading items in bulk marks the rows that match items you already have, in red where that annotator's existing labels will be replaced #65
  • See how often an evaluator agreed with the annotators, and what each annotator scored, on the evaluator run page #64
  • Rating buttons while labelling use the evaluator's own scale instead of always 1 to 5, and annotators can see the criteria values for each item #58