Writing

October 20264 min read

Accuracy stayed the same. Confidence changed.

I built Calibration Explorer to compare a classifier’s confidence with its observed accuracy. In this worked example, temperature scaling changes the probabilities while every predicted class stays the same. The app is also available on Hugging Face.

Start with your own predictions

For predictions made with 80% confidence, are about 80% correct? Calibration Explorer gives that question a practical workflow: import predictions with known outcomes, inspect confidence and accuracy, and export the measurements and settings.

I built the browser interface, strict input checks, calibration and threshold workflow, and report exports around my calibrated TypeScript library. The app does not train or run a classifier. A confidence CSV is enough to inspect calibration; temperature fitting requires the original class logits and true labels. Imported observations are processed in the browser.

Keep fitting separate from evaluation

The bundled reference is a multinomial logistic-regression classifier trained offline on the UCI handwritten-digits dataset. It is a small, reproducible example, not a model trained by the web app or evidence about modern classifiers in general.

Of the original training observations, 2,675 train the classifier, 574 fit its temperature, and 574 provide policy-validation data for assessing an acceptance threshold. All 1,797 original test observations remain outside model and temperature fitting. The saved input contains the classifier’s raw logits, class labels, and split membership.

Minimising negative log-likelihood on the calibration rows gives T = 0.72597839. Each class logit is divided by this positive temperature before softmax. The highest-scoring class remains the same, so this operation cannot improve classification accuracy by changing an answer.

Measure what changed

On the 1,797 test observations, accuracy is unchanged. Both negative log-likelihood and the chosen binned calibration measurement improve in this run.

Confidence and accuracy before and after temperature scalingUCI digits test split, 1797 observations in 15 equal-width bins. Hollow circles show original predictions; filled diamonds show scaled predictions. The diagonal marks equal confidence and accuracy. Aggregate measurements follow in the table.00252550507575100100Accuracy (%)Mean confidence (%)

OriginalScaled

Test results · 1,797 observations
MeasureOriginalScaled
Accuracy94.88%94.88%
NLL0.17600.1575
ECE0.03650.0101

Logistic regression on UCI digits. T = 0.72598, fitted on 574 separate calibration rows. Lower NLL and ECE were observed in this run; accuracy stayed the same.

ECE uses 15 equal-width bins. Small bins are noisy, and this example does not establish performance on other models or data.

Inspect the recorded measurements

NLL uses the probability assigned to the true class; lower is better. ECE averages the absolute gap between confidence and accuracy within bins, weighted by the number of observations. Its value depends on the bins and sample size. A smaller ECE here is an observed result, not proof of perfect calibration.

Follow the released workflow

The released 0.2.0 interface follows the same reference from policy validation through calibration, a locked threshold, explicit test review and report export. The numerical dependency and reference results are unchanged.

Silent step-by-step capture of the released interface, assembled from six actual screenshots. This is not a real-time recording.

Calibration Explorer 0.2.0, captured 4 October 2026 with the public UCI digits reference. The 80% threshold is illustrative. These agent-operated checks involved zero independent human participants.

Read the six steps and view the screenshots
  1. Load the reference

    Load the handwritten-digit reference. Policy validation is selected; the 1,797 test rows remain unopened.

    Calibration Explorer with the UCI handwritten-digit reference loaded, Policy validation selected and the test split unopened.
  2. Inspect a linked reliability bin

    Select original bin 15: 412 predictions, 98.3% mean confidence and 100% observed accuracy.

    Original reliability chart with bin 15 selected and linked details showing 412 predictions, 98.3% confidence and 100% accuracy.
  3. Fit on calibration rows

    Fit temperature using only 574 calibration rows. The fitted value is 0.72598.

    Calibrate stage showing a fitted temperature of 0.72598 from 574 calibration observations, with Policy validation still selected.
  4. Lock an illustrative threshold

    Lock an illustrative 80% threshold on policy validation before opening the test split.

    Decide stage with an 80% threshold locked at temperature 0.7260 on policy validation; the next action is Review test results.
  5. Review the test comparison

    Review 1,797 test predictions. Accuracy stays at 94.9%; NLL changes from 0.1760 to 0.1575.

    Test comparison showing 1,797 observations, unchanged 94.9% accuracy, ECE changing from 0.0365 to 0.0101 and NLL from 0.1760 to 0.1575.
  6. Preview and save the report

    Preview the test report. Source rows are excluded; the actual HTML download matches this preview.

    Export stage with HTML, JSON and CSV controls, source rows unchecked, and a preview labelled as the same snapshot used by the HTML download.

Confidence affects which predictions you accept

The recorded workflow uses an illustrative 80% confidence threshold, fixed before test inspection and assessed on policy-validation data. It was not optimised or recommended. Applied to the scaled test predictions, it accepts 1,635 observations with 31 errors and abstains on the remaining 162.

Those counts describe a specific decision rule on a specific dataset. They do not choose the cost of an error or an abstention for another application. Temperature can also change confidence rankings across observations, so acceptance outcomes need to be measured rather than inferred from unchanged accuracy.

Keep the record and the limits

This result was recorded with Calibration Explorer 0.1.0 and @m-sanchez/calibrated 2.0.1. The public experiment record preserves the input hash, fitted temperature, split, binning, threshold, and measured outcomes. The source and reproduction instructions include the saved predictions and data preparation.

Using the same temperature elsewhere requires the same model, preprocessing, and class order. Changing the population or encountering drift calls for fresh evaluation. Even the reference’s training-derived partitions cannot be claimed writer-independent because per-observation writer identities are unavailable.

The app’s lock records whether a policy was fixed before test inspection in that browser session. It cannot establish whether someone examined the data elsewhere, and it makes no automatic deployment decision. Open Calibration Explorer to inspect the reference or bring your own predictions.