Packaging And Registry

ONNX and Precision: Every Prediction Moved a Little, and No Decision Moved

0 of 28 complete

0%

Contents

Back|Packaging And RegistryONNX and Precision: Every Prediction Moved a Little, and No Decision Moved
1/28
62 min left
Prerequisites
Library Version Skew: Old Model Files Broke on Upgrade, and Retraining Made a Different ModelrequiredSaving Formats Compared: The Scores Never Moved, the Bytes DidrequiredContainer Images: Most of the Size Was the Base and the Libraries, and Slimming Kept Every ScorerequiredWhat a Feature Is: A Better Model or a Better Feature?required
Related Topics
Wrong Labels: How Many Can a Model Survive, and Can You Find Them?Data Engineering for MLRebalance, or Just Move the Threshold? Measured on Rare ClassesData Engineering for MLLeakage Before the Split: How Pure Noise Scored 93% AccuracyData Engineering for MLWhat a Feature Is: A Better Model or a Better Feature?Features and Feature StoresFeature Freshness: Two Weeks Stale Cost Nothing I Could MeasureFeatures and Feature Stores
1 of 28

A Photo Printed Smaller

Let me start with a photo, like the one in the picture.

Imagine you have a big, sharp photo of a lake and mountains. You print the same photo much smaller, to fit in a wallet. Almost everything is still there: the lake, the trees, the sky. If you hold the two side by side, you see the same picture.

A flat illustration of a man at a desk holding a large photo of a mountain lake in one hand and a small print of the same photo in the other, comparing them. Below the picture: converting a model to ONNX is like printing a photo smaller; almost everything survives, and two dots that sat very close together can become one.

But look closely at the small print. Two tiny dots that were separate in the big photo now sit in the same spot, because the small print has fewer dots to work with. Most of the time nobody notices. Once in a while, those two dots were the part that mattered.

A trained model can be "printed smaller" too. In this lesson I copy one real model into a format called ONNX, which stores its numbers with fewer digits. Then I check, row by row, which answers survived.

Where This Lesson Starts

This lesson uses the model from the features chapter. If it is new to you, please read what a feature is first. That lesson built six features for each customer of a real online shop. It asked one question: at the start of a month, will this customer buy something in the next 30 days? The model is a gradient boosted tree model from scikit-learn, and it scored a test AP of 0.5450. AP, average precision, is a score from 0 to 1 for how well the model ranks the buyers above the others.

Three lessons of this chapter lead here. Saving formats compared saved the model eleven ways, and every file gave back every prediction to the last bit. It ended by saying a conversion to ONNX would open that question again. Library version skew showed that a saved model often breaks when scikit-learn changes, and mentioned ONNX as a way to run a model without scikit-learn, untested. And container images found that moving from a Mac to Linux changed the very last digit of 33 predictions.

Here I test ONNX for real, on the same model and the same rows.

The Words You Need First

Please read this slide slowly if any word is new. Every slide after it uses these words.

A hand-drawn list of nine words, each in a box on the left with its meaning on the right. ONNX: a file format for models that does not belong to any one library. Converter: a tool that writes a model in another format; here skl2onnx. Runtime: the program that runs the file and gives answers; here onnxruntime. Float64: a number stored in 64 bits, about 16 correct digits; scikit-learn's default. Float32: a number stored in 32 bits, about 7 correct digits; ONNX's usual default. Threshold: the number a tree compares a value with; at most it goes left, else right. Path: the list of left and right turns a row takes down one tree. Zipmap: a converter option; off means the chances come out as a plain table. Protobuf: the library that writes the ONNX file's bytes. Below: model, test AP and the six columns mean what they meant in the features chapter.

ONNX (Open Neural Network Exchange) is a file format for models. A pickle file, from earlier lessons, can only be loaded by Python with scikit-learn installed. An ONNX file describes the model with a fixed list of standard steps, so other programs can run it.

A converter turns a scikit-learn model into an ONNX file. The one for scikit-learn is called skl2onnx. A runtime is the program that loads the file and computes answers. I use onnxruntime, which runs on a server with no scikit-learn at all.

A computer stores a number with a fixed number of binary digits, called bits. float64 uses 64 bits and keeps about 16 correct decimal digits. float32 uses 32 bits and keeps about 7. scikit-learn works in float64. ONNX converters usually write float32.

Each tree in the model asks questions like "is money at most 321.65?". The number in the question is a threshold. If the answer is yes, the row goes left; if no, it goes right. The turns one row takes down one tree are its path, and the path decides which leaf value the tree adds.

What the Sources Say

Before converting anything, I read what the tools themselves say. Every quote here was found word for word in the released source or at a fixed commit, and all of them are in results/onnx-factcheck.json.

Six quote cards in two columns, each with a logo and a source line. skl2onnx 1.20.0 docs: sklearn-onnx is using single floats by default. skl2onnx 1.20.0 docs: In most cases, float and double comparison gives the same result. skl2onnx 1.20.0 docs: A decision tree trained for a regression is not a continuous function. Therefore, even a small dx may introduce a huge discrepency. ONNX 1.23.1, TreeEnsembleClassifier: nodes_values: list of floats; output Z: tensor(float). scikit-learn 1.9.1, common.pyx and _predictor.pyx: X_DTYPE = np.float64; if data_val is at most node.num_threshold. ONNX 1.23.1: This operator is DEPRECATED. Please use TreeEnsemble with provides similar functionality. Caption: in the ONNX file the thresholds are 32-bit; scikit-learn compares in 64-bit.

skl2onnx's own documentation is open about precision. It says "sklearn-onnx is using single floats by default". Single is another name for float32, and double for float64. It explains that most scikit-learn models compute "with double, not float", and that "the conversion to float introduce small discrepencies compare to double predictions". (The spellings are theirs.)

Then it warns about trees. "In most cases, float and double comparison gives the same result. However, the probability that both comparisons give different results is not null." And: "A decision tree trained for a regression is not a continuous function. Therefore, even a small dx may introduce a huge discrepency." A continuous function changes a little when its input changes a little. A tree does not: a tiny change can send a row down the other branch.

The ONNX specification shows why. The model becomes one step called TreeEnsembleClassifier, whose thresholds are a "list of floats" and whose output is tensor(float), both float32. scikit-learn's own code keeps the input as np.float64. It goes left if data_val <= node.num_threshold, in 64 bits. So the question is: how often does that difference matter, for this model?

What Converting Does

Here is the whole trip, in the order it happens.

A sequence diagram with three lifelines: scikit-learn, skl2onnx and onnxruntime. Step 1, scikit-learn gives skl2onnx the trained model. Step 2, skl2onnx gives onnxruntime model.onnx, 121,006 bytes. Step 3, scikit-learn sends onnxruntime the rows, as float32. Step 4, onnxruntime sends back the chances, float32. Below: step 1 runs once, where scikit-learn is installed; steps 3 and 4 run on the server, which needs only onnxruntime and numpy; the file holds 3,294 tree nodes, the same 54 trees. Caption: I compared step 4's chances with scikit-learn's own, row by row.

The converter reads every tree in the model and writes every node into the file: which column it looks at, its threshold, and where to go next. The model has 54 trees with 3,294 nodes, and the file holds all 3,294.

I used the converter's option zipmap set to off. By default, skl2onnx's docs say, it "converts that matrix into a list of dictionaries where each probabily is mapped to its class id or name". A dictionary here is a small lookup table. With zipmap off, the chances come out as a plain table of numbers, one row per customer, which is easy to compare with scikit-learn's.

At prediction time the server sends the rows in, as float32 numbers, and gets the chances back. It never imports scikit-learn.

How the Lab Was Built

I wrote the design into the docstring of scripts/labs/packaging/onnx_precision.py before it first ran. Before that I had installed the libraries and read the converter's source code, which told me it writes the thresholds as float32. I had not converted this model.

A page in four labelled zones, two by two. The model: lesson 1's model, test AP 0.5450, retrained with seeds 0 to 19. Converted: skl2onnx, float32 input, zipmap off; and float64 input, if it converts. Compared: all 26,851 test rows: differences, rank in the month, decisions at 0.5 and the top fifth, test AP. Located: my own walk down the trees in float64 and float32, to see which rows take another path. Caption: guesses, written first: every row differs a little; under 1% take another path; float64 converts.

The model, twenty times. The lab trains lesson 1's model and stops if its test AP is not 0.5450. It also trains it with seeds 1 to 19. A seed is the number that fixes the random choices in training, so each seed gives a slightly different model. The features chapter's rule is to never trust one seed, and here that rule mattered.

Converted two ways. Each model is converted with float32 input, ONNX's usual default, and again with float64 input, to see if 64 bits avoids the rounding.

Compared four ways. For every test row, the lab compares onnxruntime's chance of buying with scikit-learn's. It counts the rows that differ at all and the largest difference. It counts the rows that change place in their month's ranking. It counts the decisions that flip, at a cut-off of 0.5 and at "the top fifth of each month", a rule a shop might use to pick who gets an offer. And it computes test AP, one per month, then the mean.

Located. I wrote my own small walk down every tree, in numpy, that can compare in float64 or float32. It tells me which rows take a different path, and at which split.

The First Run Failed

The first run never got to a single prediction. Both conversions stopped with an error.

Three rows, each with a logo. C1, as designed, PyPI logo: skl2onnx 1.20.0, protobuf 7.36.2; failed: TypeError: Expected an int, got a boolean. C2, protobuf 6, PyPI logo: skl2onnx 1.20.0, protobuf 6.33.6; converted; the whole lesson runs here. C3, main branch, GitHub logo: skl2onnx from GitHub, commit 449ec05, protobuf 7.36.2; converted; other predictions, a later slide. Below: all three, scikit-learn 1.9.1, onnx 1.23.1, onnxruntime 1.30.0, Python 3.13; C1's error came from protobuf, when skl2onnx wrote False into a list of whole numbers. Caption: a converter released before your libraries can break on a detail.

The message was TypeError: Field onnx.AttributeProto.ints: Expected an int, got a boolean. I read the converter's code. For every leaf of every tree, skl2onnx 1.20.0 writes the Python value False into a list that must hold whole numbers. protobuf, the library that writes the file's bytes, refused it.

skl2onnx 1.20.0 was uploaded on 30 January 2026. protobuf 7.34.0 came out on 27 February 2026, a month later, and my venv had protobuf 7.36.2. So I wrote a second design into the docstring, labelled as decided after the first run, and tried three separate venvs on the same saved model.

C2 was the same venv with protobuf 6.33.6, the newest 6.x. It converted. C3 used skl2onnx's unreleased main branch from GitHub, at commit 449ec05, where the line now reads int(nodes_missing_value_tracks_true) # cannot use boolean in onnx. It converted too.

Everything in this lesson runs in C2, the newest combination of released packages that works. C3 comes back later, because it gave different predictions.

The Lab's Report, Running

This is a real recording of the report script, onnx_report.py, on the laptop where the lab ran.

A terminal recording of onnx_report.py. It prints that the stored arrays match their sha256 and seed 0 retrained has test AP 0.5450, lesson 1's; that skl2onnx 1.20.0 with protobuf 7.36.2 gives TypeError, Expected an int, got a boolean; that with protobuf 6.33.6 it converts and the float64 file will not load; ONNX 121,006 bytes, pickle 202,700; the minimal venv gives identical output and cannot load the pickle. Seed 0, released, float32: rows differ 26,851, largest 1.94e-07, rank moved 22, flips 0 and 0, test AP 0.5449816308 to 0.5449815920, path changes 0. Seed 0, main branch: largest 0.0434, rank moved 11,316, flips 1 and 0. A table of 20 seeds with path changes, largest difference and rank moves for the released converter, and largest difference and flips for main. Then: released, 0 flips in 20 seeds, 20-seed AP difference minus 8.3e-07 with interval minus 1.7e-06 to minus 1.6e-07; main, 3.7e-06 with interval minus 3.2e-05 to 3.8e-05. Released path changes, all on money: 12346 at -64.68, 17471 at 321.65. Two lines labelled review: R1, a hand-built float64 TreeEnsemble over 20 seeds, largest 5.6e-16, 0 flips; R2, against scikit-learn on float32 rows, rows above 1e-6: main 0 per seed, released 91 to 366. No timings: 1-minute load 3.92 at the start and 8.90 at the end, rule below 2.0. Last line: all 586 checks agree with the stored lab.

The report imports nothing from the lab. It checks the sha256, a fingerprint, of every stored array. It retrains seed 0, converts it again, and runs onnxruntime again. It recomputes every comparison for all 20 seeds with its own code and scikit-learn's own AP function. It redoes both bootstraps with the same random draws. And it starts the other venvs again: C1 must still fail the same way, C3 must give the same output, and the small venv on the "What the Server Needs" slide must still give identical output. After an independent review it also reruns two added checks, R1 and R2, explained on the float64 and main-branch slides. All 586 checks agreed.

The report caught one claim of mine. After the main run I wrote down that every changed path belonged to one customer. I had read only part of the output. There were two customers, and the report's check failed until I fixed my sentence. The lab's docstring records the correction.

The Headline: Every Row Moved, No Decision Did

Here is seed 0, the chapter's own model, converted with float32 input.

Three panels, with the number 0 large in the corner over the words decisions flipped. Rows that differ: 26,851, of 26,851 test rows. Largest difference: 1.94e-7, about 2 in 10 million; mean 2.99e-8. Decisions flipped: 0, at 0.5, or at the top fifth of each month. Below: test AP: scikit-learn 0.5449816307508443, ONNX 0.544981591960395; 22 rows changed place in their month. Caption: not equal anywhere, and the same in every decision.

Every one of the 26,851 chances of buying was different from scikit-learn's. Not one row was equal to the bit. Compare that with lesson 2, where every saving format gave back all 26,851 to the last bit, and lesson 5, where 33 rows changed in the last digit.

But the differences were tiny. The largest was 1.94e-7, which means 0.000000194, about 2 in 10 million. The average was smaller still, 2.99e-8.

No decision flipped. Not at a cut-off of 0.5, and not at the top fifth of each month. Test AP moved from 0.5449816307508443 to 0.544981591960395, a change of about 4 in 100 million. Per month, AP was equal to the bit in 3 of the 5 test months.

For this model, then, ONNX gave answers that were close everywhere and exact nowhere.

Where the Tiny Difference Comes From

My guess was that most of the difference would come from shorter arithmetic, and that a few rows would take another path. For seed 0, my walk down the trees found no row at all whose path changed. So the whole difference is arithmetic.

A hand-drawn sketch. Three boxes in a row joined by arrows: 54 leaf values rounded to 32 bits; added up in 32 bits; logistic step in 32 bits. An arrow leads down to a box: the chance, 32 bits, largest gap 1.94e-7. Below: measured one step at a time: scikit-learn's own answer rounded to 32 bits moves by at most 2.98e-8; float64 paths with 32-bit sums, at most 1.42e-7; onnxruntime's raw sum against scikit-learn's, at most 1.06e-6. Caption: shorter numbers, added 54 times, give a slightly different last digit.

The model makes a prediction in two steps. It adds up one leaf value from each of its 54 trees into a raw score. Then a logistic step turns the raw score into a chance between 0 and 1. In ONNX, both steps happen in float32.

I measured the pieces one at a time. Just rounding scikit-learn's final answer to float32 moves it by at most 2.98e-8. Keeping scikit-learn's paths but adding the leaf values in float32 moves it by at most 1.42e-7. And onnxruntime's raw score differed from scikit-learn's by at most 1.06e-6, before the logistic step squeezed that down.

So every row moves, because every row's 54 numbers are added with only about 7 correct digits. A sum of 54 rounded numbers is not the rounded sum.

Twenty-Two Rows Moved Place, and Nobody Overtook

If every chance moved, did the order of customers change? For seed 0, 22 rows changed their place in their month. I looked at why.

Hand-drawn bars for seed 0, pairs of customers in the same month. Pairs that swapped order: 0. Pairs that became a tie: 24. Rows that changed place: 22. Below: two customers with different chances in 64 bits can share one 32-bit value; then the sort keeps them in row order, and a few rows below them shift by one place. Caption: nobody overtook anybody.

The lab compared every pair of customers in the same month. Not one pair swapped order. If customer A had a higher chance than customer B in scikit-learn, A still had a chance at least as high in ONNX.

What happened instead is ties. 24 pairs that had two different chances in float64 got exactly the same float32 value. When two customers tie, a sort has to put one of them first, and it keeps their order in the table. So a few rows moved by one place, with nobody overtaking anybody.

That matters for any rule that picks "the top N customers". At an exact cut-off, a tie can decide who is in. Here no decision at the top-fifth cut changed.

Seed 0 Was One of the Calm Ones

If I had stopped at seed 0, I would have written "the converter is exact to about 7 digits". Then I looked at all twenty seeds.

A dot chart of the largest difference for each of seeds 0 to 19, on a scale of powers of ten from 1e-7 to 1. Two series: released skl2onnx and unreleased main. Released: most seeds sit near 1e-7; five seeds sit between 0.001 and 0.1. Main: every seed sits between 0.01 and 0.2. Below: released, 15 seeds stayed under 2.2e-7; seeds 5, 11, 12, 14 and 18 reached 4.5e-3 to 5.07e-2, in those some rows took another path; main, 4.0e-2 to 1.1e-1 on every seed. Caption: one seed would have told me the converter is exact to 7 digits; twenty said: not always.

For 15 of the 20 seeds, the largest difference stayed under 2.2e-7, just like seed 0. For the other 5 seeds, numbers 5, 11, 12, 14 and 18, it jumped by four to five powers of ten, to between 0.0045 and 0.0507. A difference of 0.05 in a chance of buying is not a rounding detail. It is a different answer.

In each of those five seeds, my walk found 4 or 5 rows that took another path through at least one tree. In the other fifteen it found none. So there are two kinds of difference here, and they are very different in size: the arithmetic, everywhere and tiny, and the path changes, rare and large.

The second series on the chart is the unreleased main branch, which I explain later.

One Customer, Seven Trillionths

I asked this after the results, and the lab labels it as an addition: which rows changed path, and why?

A number line. Above the line on the left: threshold t, in float64, -64.68. On the right: 12346's money x, in float64, -64.67999999999302. Two arrows point from the two numbers to the same dot in the middle of the line, with ticks labelled next float32 below and next float32 above on either side. Below the dot: both become this one float32: -64.68000030517578. Two cards. Scikit-learn: x is greater than t; goes right in 7 trees; chance 0.1444. ONNX: x equals t; goes left in the same 7 trees; chance 0.0937. Below: for this customer the gap between x and t is 7.0e-12; all 86 path changes in 20 seeds were on money: this customer in four seeds, and in seed 12 customer 17471, 321.65000000000003 against 321.65. Caption: a difference of 0.05 in one customer's chance, from 7 trillionths.

The largest case was seed 11 and customer 12346, the customer who appears in every lesson of the features chapter, in July 2011. Their money feature was -64.67999999999302. One of the trees asked: "is money at most -64.68?"

In float64, -64.67999999999302 is larger than -64.68, by about 7 trillionths. So scikit-learn said no, and went right. In float32, both numbers round to the same value, -64.68000030517578. So ONNX found them equal, said yes, and went left. The same thing happened in 7 of that model's trees. scikit-learn gave a chance of buying of 0.1444; ONNX gave 0.0937.

Every path change in all 20 seeds was this kind of case, on the money column. In seeds 5, 11, 14 and 18 it was customer 12346, in all five test months. In seed 12 it was customer 17471, whose money was 321.65000000000003 against a threshold of 321.65. Rounding never sends a row the wrong way when the two numbers are far apart. It only collapses two numbers that are almost equal.

Where That Threshold Came From

Why would a tree have a threshold of exactly -64.68, so close to this customer's value? I asked after the results, and traced it for customer 12346.

Eighteen small cards in three columns, one per cutoff month, each with customer 12346's money to every digit. 2010-03: 100.0. 2010-04 to 2010-06: 127.05. 2010-07 to 2010-10: -59.18000000000001. 2010-11, 2010-12 and 2011-01: -64.68, highlighted. 2011-02 and 2011-03, then the five test months 2011-07 to 2011-11: -64.67999999999302, highlighted. Below: in 3 training months, November 2010 to January 2011, the sum was exactly -64.68: the only 3 training rows with that value, all this customer; from February 2011 a new pair of invoice lines, a purchase and its return, left the sum 7.0e-12 higher. Caption: the tree split on a value this customer once had, then later missed by 7 trillionths.

In three training months, November 2010 to January 2011, customer 12346's money was exactly -64.68. Those were the only three training rows with that value, and in seeds 5, 11, 14 and 18 the trees used -64.68 as a threshold.

Then, on 18 January 2011, the shop data has two lines for this customer: 74,215 units bought at £1.04, which is £77,183.60, and the same amount returned 16 minutes later. In exact arithmetic they cancel. In float64 they left a trace: from February 2011 on, the sum was -64.67999999999302. The customer's real money did not change; only the last bits of the sum did.

So the threshold was the customer's own old value, and the new value missed it by 7 trillionths, in a direction that only float64 can see. I did not trace customer 17471's history.

The lab also counted this for the seed-0 model. In five of the six columns, every single threshold equals some training value. Frequency is the exception, with thresholds halfway between whole numbers. Test values that land exactly on a threshold are common: 16,965 for return share and 16,675 for products. When both numbers are equal in float64 they are also equal in float32, so with the released converter those rows are safe. The danger is a value a hair above a threshold.

Does It Matter?

So far: every chance moves a little, and in a quarter of the seeds a few rows move a lot. Did any of it change what the model is used for?

Three panels. Decisions flipped: 0, in all 20 seeds, at 0.5 and at the top fifth. Rows that moved place: 4 to 4,337, per seed, of 26,851. 20-seed AP gap: -8.3e-7, 95%: -1.7e-6 to -1.6e-7. Below: the AP gap is scikit-learn minus ONNX, over 1000 resamples of customers; it does not cross zero, so by the chapter's rule it is measurable: ONNX scored higher, by less than two millionths. Caption: measurable, and far too small to change a choice.

No decision flipped, in any of the 20 seeds, at 0.5 or at the top fifth. The large path changes all happened to customers whose chance of buying was far from either cut-off. Customer 12346, for example, went from 0.1444 to 0.0937, both well below 0.5 and outside the top fifth.

The ranking moved more in the five unusual seeds: up to 4,337 rows changed place in seed 11, because one customer moving down by 0.05 shifts everybody they pass by one place.

Test AP barely moved. The chapter's rule for a question like "is it different?" is a paired bootstrap of the 20-seed mean: draw the test customers again at random, 1,000 times, within each month, and recompute both scores on each draw. The 95% interval of scikit-learn minus ONNX was -1.7e-6 to -1.6e-7. It does not cross zero, so the difference is measurable. ONNX scored higher on average, by less than two millionths.

So, for this model and these rows: measurable, and much too small to change any choice I would make. But one customer's chance moving by 0.05 is a real change for that customer, and in another model, near a cut-off, it would be a flipped decision.

Float64 Through This Converter Was Not an Escape

If 32 bits is the problem, why not ask for 64? skl2onnx 1.20.0's classifier converter accepts an input declared as float64. So I tried it on all 20 seeds.

A page in two parts, titled skl2onnx 1.20.0's 64-bit option gave a file that would not load. Top: the converter wrote it, 121,006 bytes, and onnx.checker passed; onnxruntime said, in a box: Type Error: Type (tensor(double)) of output arg (probabilities) of node (TreeEnsembleClassifier) does not match expected type (tensor(float)). Then: the file keeps 32-bit thresholds and a 32-bit output, and says the output is 64-bit. Bottom, headed my simulation, not onnxruntime, 64-bit values against the file's 32-bit thresholds: 19 to 118 rows per seed take another path; 2 and 8, most flips in a seed, at 0.5 and the top fifth; 164 of 164 seed 0 splits where x equals t exactly. Then: after the review, by hand: the newer opset-5 TreeEnsemble in float64 ran in onnxruntime, largest gap 5.6e-16 over 20 seeds; the format can do 64 bits; this converter does not write it. Caption: with this converter, rounding only the thresholds is worse than rounding both sides.

The converter wrote a file for every seed, and the ONNX checker passed it. But onnxruntime refused to load any of them: "Type (tensor(double)) of output arg (probabilities) of node (TreeEnsembleClassifier) does not match expected type (tensor(float))". The file says its output is float64, while this tree step, by the ONNX specification, always outputs float32. Its thresholds were still stored as float32 too. This converter always writes that older step. Its code fixes op_version=1 for this model, and keeps the thresholds in float32 unless the op version is 3 or more.

What would 64-bit input against those 32-bit thresholds have done? onnxruntime would not tell me, so I measured it with my own tree walk, and I label it as a simulation. It was worse than float32. Between 19 and 118 rows per seed took another path, against 0 to 5 with float32. In four seeds a decision at 0.5 flipped, up to 2 in one seed, and one seed flipped 8 at the top fifth.

Here is the reason, and it surprised me. When both the value and the threshold are rounded to float32, a value exactly equal to a threshold stays equal, and goes left, just like in scikit-learn. When only the threshold is rounded, the rounded threshold can land just below the value, and an exactly equal value goes right. For seed 0, all 164 splits where the simulation disagreed had a value exactly equal to the threshold.

After an independent review, I checked whether the format itself can do 64 bits. It can. ONNX's newer step, in version 5 of the operator set, accepts float64 numbers. I built such a file by hand from each model's own trees, with float64 thresholds and leaf values, and onnxruntime 1.30.0 ran it. Over all 20 seeds the largest difference from scikit-learn was 5.6e-16, test AP was equal to the bit in every month, and no decision flipped. Customer 12346 at seed 11 got 0.14439673196282055, against scikit-learn's 0.14439673196282057.

The Unreleased Fix: Right for Float32 Data, Not for These Rows

skl2onnx's main branch on GitHub, C3 from the first-run slide, has a change for exactly this model type: pull request #1227, merged on 25 February 2026 and not yet in a release. Its comment says the float64 thresholds "must be adjusted so that float32 comparisons in ONNX give the same routing decisions as sklearn's float64 comparisons."

A bar chart titled on these float64 rows, the unreleased fix moved more, with the number 120 large in the corner over the words rows, seed 0, main. Two bars per seed, 0 to 19: rows whose path changed with the released 1.20.0, and with main at commit 449ec05. Released bars are 0 for most seeds and 4 or 5 for five seeds. Main bars are between about 90 and about 370 on every seed. Below: main lowers 548 of seed 0's 1,620 thresholds to the 32-bit number below them; in seed 0, all 156 first splits it changed had x equal to t exactly, and it sends them right; against scikit-learn on rows cast to float32, main had 0 rows off by more than 1e-6 in all 20 seeds; customer 15974 went from 0.5021 to 0.4588: a flip at 0.5. Caption: main matches scikit-learn on float32 data; these rows were float64.

On my test rows it did the opposite. With main, between 91 and 366 rows per seed took another path, against 0 to 5 with the release. In seed 0, 120 rows changed, the largest by 0.0434, and one decision at 0.5 flipped: customer 15974 in November 2011 went from 0.5021 to 0.4588. Over 20 seeds, main flipped up to 3 decisions at 0.5 and up to 16 at the top fifth in one seed.

I checked what changed in the file. Main lowers 548 of seed 0's 1,620 thresholds: every one whose float32 rounding landed above the true value. The pull request says why: scikit-learn "compares float32 inputs against float64 thresholds". That holds when the data reach scikit-learn as float32. My test rows are float64, and this model compares them in float64, as its _predictor.pyx shows. So a value exactly equal to a threshold goes left in scikit-learn and right past the lowered threshold. In seed 0, all 156 changed splits were exact equalities.

After the review I ran scikit-learn again on the same rows cast to float32 first. Against that, main had 0 rows off by more than 1e-6 in all 20 seeds, and its largest gap was 2.2e-07. The release had 91 to 366 rows off by more than 1e-6, up to 0.11. So main matches scikit-learn when the data reach it as float32, and moves rows when they are float64 and a value sits exactly on a threshold. Neither converter is simply better; which one matches depends on the dtype of your data.

Two honest limits. This is unreleased code at one commit, and it may change before a release. And on test AP, main's 20-seed interval, -3.2e-5 to 3.8e-5, crosses zero: by the chapter's rule it cannot be told apart from customer luck. The harm here is to single rows, not to the average.

Even the Thread Count Moved the Last Digits

Two smaller things turned up after the results, and both are labelled additions in the lab.

Hand-drawn bars of rows that differ from the default run of the same file, seed 0: 1 thread, 5,731; 2 threads, 3,184; 4 threads, 3,145. Below: by at most 3.0e-7, and no decision flipped; and converting the same model twice gave different bytes: the graph gets a random name, b3e5ee8e then 6ffddc83, while the trees inside were identical. Caption: compare outputs with a tolerance, and compare files by their contents, not their bytes.

Threads. A thread is one worker inside a program; onnxruntime splits the rows across several. I ran the same seed-0 file with 1, 2 and 4 threads. With one thread, 5,731 rows differed from the default run, by at most 3.0e-7. Adding float32 numbers in a different order gives different last digits. The student demo shows the same thing: one row predicted alone gave a slightly different chance than the same row inside the full batch.

Bytes. Converting the same model twice gave two different files. The only difference was the graph's name, a random id such as b3e5ee8ed9b346a5be29a1433449814e. The trees inside were identical. Reproducible artifacts, lesson 6 of this chapter, asks exactly this question for pickle files.

Both point to one habit. Compare predictions with a small tolerance, never with "equal". And compare model files by what they contain, not by their bytes.

What the Server Needs

The reason to use ONNX at all is the server side. What does loading the file actually need?

An isometric drawing of two stacks. Left, full venv, 447.7 MB, a lower block for onnxruntime and a tall upper block. Right, onnxruntime plus numpy, 105.1 MB, a lower block for onnxruntime and a short upper block. Below: lower block, onnxruntime itself, 78.1 MB in both; the full venv's biggest folders: pyarrow 126.3, scipy 82.7, onnxruntime 78.1, pandas 49.2, sklearn 38.6; the ONNX file was 121,006 bytes, the pickle 202,700. Caption: in the small venv the pickle failed, ModuleNotFoundError: No module named 'sklearn'; the ONNX file gave identical output.

I made a second, minimal venv with only onnxruntime 1.30.0 and numpy 2.5.3. The installer added three small helpers that onnxruntime needs: flatbuffers, packaging and protobuf. In that venv, scikit-learn, scipy, pandas, skl2onnx, onnx and joblib all fail to import.

The ONNX file loaded there and gave identical output, to the bit, to onnxruntime in the full venv. The pickle of the same model did not load at all: ModuleNotFoundError: No module named 'sklearn'.

The minimal venv's packages took 105.1 MB on disk, against 447.7 MB for the full lab venv, though the full one also holds pandas and pyarrow, which a server may not need. onnxruntime itself is 78.1 MB of it. The ONNX file was 121,006 bytes, 0.60 times the 202,700-byte pickle. I guessed it would be bigger.

This is ONNX's real promise, and it held: the server no longer depends on scikit-learn's version, which is the whole problem of lesson 3.

My Guesses Before the Run, Checked

I wrote seven guesses into the lab before it ran. Here they are against the results.

  1. "skl2onnx 1.20.0 converts this scikit-learn 1.9.1 model without error." Wrong. It failed with protobuf 7, and worked with protobuf 6.

  2. "f32: almost every row differs from scikit-learn by a tiny amount, max abs difference below 1e-6 on rows whose path did not change; some rows change path because x32 == t32, and those differ by more, up to about 0.05. I guess fewer than 1% of rows change path." Right. Every row differed, by at most 2.2e-7 where the path did not change; at most 5 rows changed path, by up to 0.0507.

  3. "f32: test AP equal to 4 decimal places, not to the bit." Right.

  4. "f32: at most 5 decisions flip at 0.5, at most 20 at the top-fifth cut; more than 100 rows change rank within their month (mostly new ties among float32 values)." Right on flips: none. Wrong on rank for seed 0: 22 rows. Right that ties, not swaps, caused it.

  5. "f64: the converter accepts float64 input, but the thresholds are still float32 in the file, so some rows still change path (fewer than f32), and the output is still float32." Half right: it converted with float32 thresholds, but onnxruntime would not load it, and my simulation found more path changes, not fewer.

  6. "the ONNX file is bigger than the pickle, between 1.5x and 3x." Wrong: 0.60 times.

  7. "the minimal venv predicts bit-identical to the full venv's onnxruntime, and cannot load the pickle." Right.

Try It Yourself

The full lab trains twenty models, uses four venvs and runs a bootstrap. I wrote a small demo that shows the core of it with one Python.

A page in three labelled zones side by side, headed onnx_demo.py, designed before it ran. Train: seed 0 and seed 11, lesson 1's six columns. Convert: float32, zipmap off, run with onnxruntime; then float64. Compare: every row, the ranks, two decisions, test AP; customer 12346. Caption: it printed 26,851 rows differ, largest 1.94e-7, 0 and 0 flips; float64 would not load; 12346: 0.1444 against 0.0937.

I wrote the demo's design into its docstring after the lab had run and before the demo first ran. After its first run I changed three small things, and the docstring says so. The one that matters: it had predicted customer 12346's row on its own and got 0.09372258186340332, while the lab, predicting all rows at once, got 0.09372246265411377. The demo now predicts all rows and prints both.

A real screenshot of VS Code with onnx_demo.py open at the top of the file, showing its docstring: what it needs, the pip install line with protobuf below 7, how to run it, and the design written before it first ran.

Before you run this lab. You need Python 3 with pandas, pyarrow, scikit-learn, skl2onnx, onnx and onnxruntime. These are the versions I used, with the one pin that matters: pip install scikit-learn==1.9.1 skl2onnx==1.20.0 onnx==1.23.1 onnxruntime==1.30.0 "protobuf<7" pandas pyarrow openpyxl. I ran it inside the lab's venv, which you can make the same way and activate with source ~/lab-data/venv-pkg-onnx-pb6/bin/activate. First run python fetch_data.py from the scripts/labs/features folder; it downloads the shop data once, about 46 MB.

Then, inside the examples folder, run python onnx_demo.py. I ran it on a Mac with an Apple processor. These libraries run on Windows and Linux too, but I have not checked the numbers there; the last digits may differ, as lesson 5 found for Linux.

Round a Number Yourself

This box holds the real results from the lab for all 20 seeds. It needs nothing but Python, so it runs in your browser. It loads no model.

Press Run. It rounds customer 12346's money and the threshold to float32 the way ONNX stores them, and shows which way each side goes. Then it prints seed 11's results. Change X, T or SEED and run again. Try X = -64.6799 to see a gap that float32 can still see.

The report script writes this box from the lab's stored results and checks what it prints. The function f32 uses Python's struct module to store a number in 4 bytes and read it back, which is exactly the rounding ONNX does.

The Lab's Code, Piece by Piece

The lab is one file, scripts/labs/packaging/onnx_precision.py. It runs in the C2 venv.

data imports lesson 1's own feature code and builds the train and test rows. train fits the model for one seed. convert calls skl2onnx with float32 or float64 input and zipmap off, and run_ort runs a file with onnxruntime.

compare takes scikit-learn's chances and another set of chances. It returns everything on the headline slides: rows that differ, the largest and mean difference, rank changes, flips at both cut-offs, and AP per month. ap_one computes AP in exactly scikit-learn's order of operations, and the lab checks it gives the same bits as scikit-learn's function.

tree_tables and walk are my own tree walk. They read each tree's nodes and send every row down it, comparing in float64 or float32. adj32 copies the main branch's threshold rule, so the walk can predict which rows main moves. first_disagreements finds, for each changed path, the first split where the two comparisons disagree, and records whether the value equalled the threshold.

How to Ship a Model as ONNX

Here is how I would convert a model like this one, using only what this lab measured.

A flowchart. A model to serve without scikit-learn leads to: convert with a released skl2onnx, float32. Then: run stored rows, compare chances, ranks, decisions. Then a question: a decision flipped, or a big gap? No leads to: ship the .onnx with the rows and their expected chances. Yes leads to: find the row and its threshold, keep scikit-learn. Ship leads to: compare again after every converter or runtime upgrade. Below: compare with a tolerance, never equal bits: every row differed here, by design of the format. Caption: run the comparison on many seeds or many rows: one seed hid the real case.

The chart turns on one question, asked after the comparison: did any decision flip, or is one row's gap far bigger than all the others? Here are the same steps in words, each with the number from this lab behind it.

  1. Pin the whole set. skl2onnx, onnx, onnxruntime and protobuf, together. Here a protobuf release a month after skl2onnx's broke the conversion. Lesson 4 of this chapter shows how to lock versions.

  2. Use a released converter, and match the dtype. Convert with float32 input, its default; with skl2onnx 1.20.0's classifier converter, float64 input gave a file that would not load. If your rows are float64, the release matched scikit-learn better here; if they are float32, main's change targets your case. One option I did not test: cast your data to float32 at the source, before both training and serving, so every tool sees the same numbers.

  3. Keep test rows and scikit-learn's chances next to the file. After converting, run onnxruntime on those rows and compare three things: the largest difference, the order of customers, and every decision your system makes from the chance.

  4. Look at the worst row. If one row differs by far more than the rest, find its path. A value a hair above a threshold is the usual cause here.

  5. Test more than one model. Seed 0 had no path change at all; 5 of 20 seeds did. If you retrain often, compare every new model, not just the first.

When to Use ONNX, and When Not To

Use ONNX when the server should not depend on scikit-learn. Here the file ran in a venv a quarter the size of the full one, and it did not care which scikit-learn trained it. That removes the whole problem of lesson 3.

Use it when tiny differences are fine for every decision you make. Here no decision flipped in 20 seeds with the released converter.

Do not use it blindly for a model whose decisions sit near a cut-off. One customer's chance moved by 0.05 here. Had it been near 0.5, it would have flipped.

Do not expect skl2onnx's float64 option to fix precision. With its 1.20.0 classifier converter it gave a file that would not load. The format can do 64 bits, through the opset-5 TreeEnsemble that onnxruntime runs, but this converter does not write it.

The right converter depends on what dtype your data are in. On these float64 rows the release moved 0 to 5 rows per seed and the unreleased main branch 91 to 366. On the same rows cast to float32, main matched scikit-learn in every seed and the release did not.

Do not compare ONNX output with scikit-learn's using equality. Every row differed, by design of the format.

What This Lab Cannot Tell You

Two columns titled shows and cannot show. Shows: one model type, 20 seeds, one Mac; every row differs in the last digits; no decision flipped with the released converter. Which rows take another path, and why: a value that 32 bits cannot tell from a threshold. That skl2onnx 1.20.0's float64 file did not load; a hand-built opset-5 float64 file ran. Cannot show: speed, the machine was busy, so no seconds are given. Other models, other converters, or a GPU runtime. That the main branch stays this way: it is unreleased code at one commit.

One model type. Every number here is for the chapter's gradient boosted model. A model with more trees adds more rounded numbers, and skl2onnx's docs say the gap "grows with the number of trees in the forest". A linear model or a neural network would behave differently.

One Mac, one CPU runtime. onnxruntime has other ways to run a model, such as on a graphics card. I did not test them.

No timings. The load average was 3.92 when the lab started and 8.90 when it ended, because other work was running. Speed is a common reason to use onnxruntime, and I measured none.

Released code and one commit. C2 is released code. C3 is GitHub main on one day; a released version may behave differently.

Float64 was tested through skl2onnx 1.20.0's converter. Its file would not load, so those float64 numbers come from my own tree walk. After the review I also built a float64 opset-5 TreeEnsemble by hand, which onnxruntime ran to within 5.6e-16. I did not test any converter that writes that step.

Labelled additions. Five questions were asked after the main results: the changed rows, the threads, the bytes, how often values sit on thresholds, and customer 12346's history. The three-venv test was designed after the first run failed. All are labelled in the lab, the report and here, along with the one claim the report corrected.

What to Do on Monday

A hand-drawn grid of six cards, titled five checks. 1, pin the trio: converter, onnx, onnxruntime, and protobuf too. 2, keep test rows: and scikit-learn's chances on them, next to the file. 3, compare three ways: largest gap, rank in the month, decisions. 4, look at the worst row: is a value sitting on a threshold? 5, more than one seed: seed 0 had no path change; 5 of 20 did. The reason: 0.1444 became 0.0937 for one customer, from 7 trillionths. Caption: a converted model is a new model. Test it like one.

If you take one thing to work on Monday, find a model you already serve as ONNX, or plan to. Run it next to the original on a few thousand stored rows and print the largest difference. If it is around 1e-7, the arithmetic is the only difference. If it is much bigger, find that row and the split where its path changes.

Then write the comparison into your build, with a tolerance, so it runs every time the model, the converter or the runtime changes.

A closing card titled close everywhere, exact nowhere. Three numbers side by side in large type: 1.94e-7, largest difference for seed 0, on all 26,851 rows, which all differed; 0 of 20, seeds with a flipped decision, with the released converter; 9, rows of two customers took another path, in 5 of 20 seeds.

The one idea to keep: a model converted to ONNX is a new model that is close to the old one everywhere and equal to it nowhere. Here the closeness was enough for every decision, but a value seven trillionths above a threshold moved one customer by 0.05. Test the converted model on stored rows, on more than one seed, before it replaces the original.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

For seed 0, every one of the 26,851 ONNX chances differed from scikit-learn's, yet no row took another path. Where did the difference come from?

Q2

In seed 11, customer 12346's chance went from 0.1444 in scikit-learn to 0.0937 in ONNX. Why?

Q3

What happened when skl2onnx 1.20.0 converted the model with float64 input?

Q4

The unreleased main branch lowers some thresholds to fix precision. What did it do on these float64 test rows?

TreeEnsemble
ai.onnx.ml

So the format and the runtime can do 64 bits; skl2onnx 1.20.0's converter for this model does not write that step. skl2onnx's open issue #1074 reports the same load error. My hand-built file is a check, labelled as added after the review, not a converter I would ship without its own tests.

r"""Convert the chapter's model to ONNX, run it with onnxruntime, and compare with scikit-learn.

Lesson 9 of 'Packaging, Registry and Versioning'. It needs a Python with
scikit-learn, pandas, pyarrow, skl2onnx, onnx and onnxruntime, and the shop
data from the features chapter: run scripts/labs/features/fetch_data.py once
first. The versions I used, and one pin that matters (see the lesson):
    pip install scikit-learn==1.9.1 skl2onnx==1.20.0 onnx==1.23.1 \
        onnxruntime==1.30.0 "protobuf<7" pandas pyarrow openpyxl
With protobuf 7, skl2onnx 1.20.0 stops with "Expected an int, got a
boolean"; that is one of the lesson's findings. Then, inside this folder:
    python onnx_demo.py              # print the results
    python onnx_demo.py out.json     # and save them
It prints no timings.

Design, written 2026-10-02 after the lab (onnx_precision.py) had run and
before this file first ran:
  1. Train lesson 1's model (seed 0) and predict the 26,851 test rows.
  2. Convert it with skl2onnx: float32 input, zipmap off, so onnxruntime
     returns a plain table of probabilities.
  3. Run onnxruntime on the same rows, cast to float32, and compare with
     scikit-learn: rows that differ, the largest difference, rows that
     change rank in their month, decisions that flip at 0.5 and at the top
     fifth of each month, and test AP.
  4. Try float64 input: convert, then try to load it in onnxruntime.
  5. Train seed 11 too and look at customer 12346, whose path through the
     trees changed in the lab.
  onnx_report.py checks this demo's saved run against the lab.

Changed after its first run (2026-10-02): step 5 first predicted customer
12346's row on its own, and onnxruntime gave 0.09372258186340332; the lab
predicts all rows at once and got 0.09372246265411377. Step 5 now predicts
all rows, as the lab does, and prints both. One possible reason, not traced:
onnxruntime adds the trees up in another order for one row than for many
(the lab found the thread count changes the last digits the same way).
The money value is now printed as a plain float, and the float64 error
line is no longer cut short. Step 5's design line first called 12346 "the
one customer" whose path changed; the lab found two (12346 and 17471), so
those words were removed. What the demo prints did not change.

Author: Roni Das
Created: 2026-10-02
"""
import json
import struct
import sys
from pathlib import Path

import numpy as np
import onnxruntime as ort
from skl2onnx import convert_sklearn
from skl2onnx.common.data_types import DoubleTensorType, FloatTensorType
from sklearn.ensemble import HistGradientBoostingClassifier as HGB
from sklearn.metrics import average_precision_score

sys.path.insert(0, str(Path(__file__).resolve().parents[2] / "features"))
import task  # noqa: E402
from what_a_feature_is import HAND_COLS as COLS, joined  # noqa: E402

ev = task.load_events()
lab_tr, _, lab_te = task.splits(ev)
tr = joined(ev, lab_tr, task.TRAIN_CUTOFFS)
te = joined(ev, lab_te, task.TEST_CUTOFFS)
x_tr, x_te = tr[COLS].to_numpy(float), te[COLS].to_numpy(float)
y, cut = te["label"].to_numpy(), te["cutoff"].to_numpy()


def ap(p):
    """Test AP: the plain mean of the five per-month APs, as in the features chapter."""
    return float(np.mean([average_precision_score(y[cut == c], p[cut == c]) for c in np.unique(cut)]))


def ranks(p):
    """Each row's place within its own month, highest chance first."""
    r = np.empty(len(p), int)
    for c in np.unique(cut):
        idx = np.flatnonzero(cut == c)
        r[idx[np.argsort(-p[idx], kind="stable")]] = np.arange(len(idx))
    return r


def top_fifth(p):
    flag = np.zeros(len(p), bool)
    for c in np.unique(cut):
        idx = np.flatnonzero(cut == c)
        flag[idx[np.argsort(-p[idx], kind="stable")][: -(-len(idx) // 5)]] = True
    return flag


def to_onnx(model, tensor_type):
    return convert_sklearn(model, initial_types=[("X", tensor_type([None, 6]))],
                           options={id(model): {"zipmap": False}})


def onnx_predict(onx, x):
    sess = ort.InferenceSession(onx.SerializeToString(), providers=["CPUExecutionProvider"])
    return sess.run(None, {"X": x})[1][:, 1]          # outputs: label, probabilities


out = {}
# 1-3: seed 0, float32
model = HGB(random_state=0).fit(x_tr, tr["label"])
sk = model.predict_proba(x_te)[:, 1]
onx = to_onnx(model, FloatTensorType)
ox = onnx_predict(onx, x_te.astype(np.float32)).astype(np.float64)
d = np.abs(ox - sk)
s0 = {"rows_differ": int((d > 0).sum()), "max_abs": float(d.max()),
      "rank_moved": int((ranks(sk) != ranks(ox)).sum()),
      "flips_0_5": int(((sk >= 0.5) != (ox >= 0.5)).sum()),
      "flips_top20": int((top_fifth(sk) != top_fifth(ox)).sum()),
      "ap_sklearn": ap(sk), "ap_onnx": ap(ox), "onnx_bytes": len(onx.SerializeToString())}
out["seed0"] = s0
print(f"seed 0: {model.n_iter_} trees; ONNX file {s0['onnx_bytes']:,} bytes")
print(f"rows that differ: {s0['rows_differ']:,} of {len(sk):,}; largest difference {s0['max_abs']:.3g}")
print(f"rows that changed rank in their month: {s0['rank_moved']}")
print(f"decisions that flipped: {s0['flips_0_5']} at 0.5, {s0['flips_top20']} at the top fifth")
print(f"test AP: scikit-learn {s0['ap_sklearn']!r}")
print(f"         onnxruntime  {s0['ap_onnx']!r}")

# 4: float64 input
onx64 = to_onnx(model, DoubleTensorType)
try:
    onnx_predict(onx64, x_te)
    out["f64"] = {"loaded": True}
    print("\nfloat64: converted and loaded")
except Exception as e:
    msg = str(e)
    out["f64"] = {"loaded": False, "error": msg[:300]}
    print("\nfloat64: converted, but onnxruntime would not load it:")
    print("  " + msg[msg.find("Type Error"):msg.find("TreeEnsembleClassifier)") + 23])
    print("  " + msg[msg.find("does not match"):].strip())

# 5: seed 11, customer 12346
m11 = HGB(random_state=11).fit(x_tr, tr["label"])
i = int(np.flatnonzero((te["customer_id"] == 12346).to_numpy())[0])
onx11 = to_onnx(m11, FloatTensorType)
sk11 = float(m11.predict_proba(x_te)[i, 1])
ox11 = float(onnx_predict(onx11, x_te.astype(np.float32))[i])
ox11_alone = float(onnx_predict(onx11, x_te[i:i + 1].astype(np.float32))[0])
money = float(x_te[i, COLS.index("money")])
f32 = struct.unpack("f", struct.pack("f", money))[0]
out["seed11_12346"] = {"cutoff": str(te["cutoff"].iloc[i].date()), "money": repr(float(money)),
                       "sklearn": sk11, "onnx": ox11, "onnx_row_alone": ox11_alone}
print(f"\nseed 11, customer 12346 at {out['seed11_12346']['cutoff']}: money {money!r}")
print(f"  as float32 {f32!r}; -64.68 as float32 {struct.unpack('f', struct.pack('f', -64.68))[0]!r}")
print(f"  scikit-learn {sk11!r}")
print(f"  onnxruntime  {ox11!r} (all rows at once), {ox11_alone!r} (this row alone)")
if len(sys.argv) > 1:
    json.dump(out, open(sys.argv[1], "w"), indent=1)

This is a real run in VS Code's terminal, inside the examples folder.

A real screenshot of VS Code's terminal after running python onnx_demo.py. It prints seed 0: 54 trees, ONNX file 121,006 bytes; rows that differ 26,851 of 26,851, largest difference 1.94e-07; 22 rows changed rank; 0 decisions flipped at 0.5 and 0 at the top fifth; test AP scikit-learn 0.5449816307508443 and onnxruntime 0.544981591960395. Then float64: converted, but onnxruntime would not load it, with the Type Error. Then seed 11, customer 12346 at 2011-07-01, money -64.67999999999302, as float32 -64.68000030517578, the same as -64.68 as float32; scikit-learn 0.14439673196282057, onnxruntime 0.09372246265411377 for all rows at once and 0.09372258186340332 for this row alone.

When I ran it, it gave the lab's numbers for seed 0, the same float64 error, and the lab's chance for customer 12346 in seed 11. The report script checks the demo's saved run against the lab.

compat and probe ran the three-venv test after the first run failed. after holds the five questions asked after the results: the changed rows, the threads, the bytes, the thresholds, and customer 12346's history. minimal runs onnx_files/rt_only.py in the venv with only onnxruntime and numpy.

onnx_report.py checks all of it from the outside, as the recording showed.

  • Compare with a tolerance. A check for exact equality fails on every row, every time.