Skip to content

Win every scaled benchmark row - #506

Merged
godofecht merged 5 commits into
mainfrom
scale-parity
Sep 25, 2026
Merged

godofecht merged 5 commits into
mainfrom
scale-parity

Conversation

@godofecht

Copy link
Copy Markdown
Owner

The 19 canonical rows all fit in 1797 samples. The scaled matrix runs five estimators at 100, 1000 and 10000 rows by 8 and 32 features, and four of its 34 rows were losses, the worst at 0.24x. All four were algorithmic.

Coordinate descent never terminated early. The stopping test compared the epoch's largest coefficient change against 1e-16, an exact-fixpoint condition that floating point does not reach, so every fit ran its full 1000 epochs where sklearn's cd_fast takes three to ten. Replaced with sklearn's relative criterion, d_w_max / w_max < tol. Lasso at 10000x32 goes 0.24x to 2.50x and diabetes 2.60x to 7.36x. Coefficients move closer to sklearn's. The diabetes parity residual drops from 1.168e-05 to 7.809e-06.

LinearRegression solved every design by Householder QR. Tall well-conditioned designs now go through the normal equations with dgemm and dpotrf, guarded by a diagonal-ratio check that falls back to QR when the Gram matrix is near singular. 10000x32 goes 0.76x to 5.35x, diabetes to 10.21x.

The RBF SMO built the full n-by-n kernel up front. Rows are now computed on demand through one sgemv and cached. KernelSVC_RBF at 1000x8 goes 0.51x to 1.36x.

The Lasso Gram build is routed through dgemm and dgemv above a size product of 65536, with the scalar path kept below it.

Results

before after
scaled matrix 30 wins, 4 losses 34 wins, 0 losses
canonical 19/19 19/19
suite 114/114 114/114
validators green green

Scaled numbers measured locally on an Apple M4 Max at load average 5, three repeats aggregated by median, with the sklearn baseline regenerated in the same session. CI regenerates the same matrix and reports it without gating.

PCA/iris is the only other parity row that moved, 6e-08 to 0. The other 17 are unchanged.

tests/test_opt_lasso_fit.flow carried its own copy of the old stopping rule, so it was comparing two different algorithms. Its reference now uses the relative criterion, with a comment saying why.

benchmarks/scaled_flow_baseline.json is left alone. It is a Flow-only self-regression gate from a GitHub Actions artifact, it is now 1.16x to 168x behind the current code, and refreshing it from this laptop would make a fast machine the standard CI has to meet. Refresh it from a CI artifact of this run instead. README says so.

The 19 canonical rows all fit in 1797 samples. The scaled matrix runs five
estimators at 100, 1000 and 10000 rows by 8 and 32 features, and four of its
34 rows were losses, the worst at 0.24x. All four were algorithmic rather than
constant-factor.

Coordinate descent never terminated early. The stopping test compared the
epoch's largest coefficient change against 1e-16, an exact-fixpoint condition
that floating point does not reach, so every fit ran its full 1000 epochs where
sklearn's cd_fast takes three to ten. Replaced with sklearn's relative
criterion, d_w_max / w_max < tol. Lasso at 10000x32 goes 0.24x to 2.50x and
diabetes 2.60x to 7.36x. Coefficients move closer to sklearn's. The
diabetes parity residual drops from 1.168e-05 to 7.809e-06.

LinearRegression solved every design by Householder QR. Tall well-conditioned
designs now go through the normal equations with dgemm and dpotrf, guarded by a
diagonal-ratio check that falls back to QR when the Gram matrix is near
singular. 10000x32 goes 0.76x to 5.35x, diabetes to 10.21x.

The RBF SMO built the full n-by-n kernel up front. Rows are now computed on
demand through one sgemv and cached. KernelSVC_RBF at 1000x8 goes 0.51x to
1.36x.

The Lasso Gram build is routed through dgemm and dgemv above a size product of
65536, with the scalar path kept below it.

Scaled: 34 wins, 0 losses, measured on a quiet machine at load 5. Canonical
stays 19/19. Suite 114/114, all 11 validators green. PCA on iris is the only
other parity row that moved, 6e-08 to 0.

tests/test_opt_lasso_fit.flow carried its own copy of the old stopping rule, so
it was comparing two different algorithms. Its reference now uses the relative
criterion, with a comment saying why.
SMO asks the kernel cache for about a hundred rows per fit, and each row was
one cblas_sgemv. At 800 samples and 8 features a row is 6400 flops, which is
under the point where a threaded sgemv earns the cost of handing work to its
threads.

An unrolled scalar dot product below a size product of 65536 takes the whole
fit from 0.96 ms to 0.37 ms at 8 features and from 1.04 ms to 0.63 ms at 32,
best of seven on an Apple M4 Max at load 5. Above the gate the row is large
enough for BLAS to pay and nothing changes: digits at 1437 samples and 64
features is 91968 and stays on the sgemv path.

Two other shapes were tried and dropped. Folding the label product into the
row build saves a second pass and an allocation per row, worth 9 percent on a
binary fit, but it requires each one-vs-one pair to build its own cache, and
that costs 3.2x on ten classes at digits scale: 18.0 ms to 58.4 ms. Sharing
kernel rows across pairs is worth more than the fused pass. A plain scalar
loop without the four accumulators is slower than sgemv at 32 features.

Both KernelSVC_RBF rows of the scaled matrix lose on the CI runner, at 0.68x
and 0.88x, while winning on this machine. This is the first change aimed at
that gap.
A one-vs-one fit kept the whole training matrix and a coefficient matrix with
one row per training sample, so every decision value evaluated the RBF kernel
against all of it and multiplied most of the result by zero. Most rows carry a
zero dual coefficient in every pair.

The fit now compacts to the rows some pair gave a nonzero coefficient, which
is what libsvm stores. Predict on 200 test rows against 800 training samples
goes from 0.58 ms to 0.10 ms, and the whole scaled row from 1.26 ms to 0.83 ms
at 8 features and 1.54 ms to 1.18 ms at 32, best of seven on an Apple M4 Max
at load 2. Against scikit-learn's 1.37 ms and 1.78 ms on the same machine that
is 1.65x and 1.51x, from 0.93x and 1.16x.

A fit that finds no support vector keeps one zeroed row, so predict still has
a shape BLAS takes and every decision value comes back as its bias.
test_opt_svc_predict asserted that a fitted KernelSVCMulti records the whole
training set, which no longer holds now that the fit compacts to support
vectors. It now checks the model keeps between one row and the training set
size.

The assertion that matters is unchanged: the one-vs-one decision values still
match the per-pair reference path, to 1.1e-06 on iris.
benchmarks/scaled_flow_baseline.json dated from GitHub Actions run 32002315336
and every row had moved since, most by an order of magnitude: KernelSVC_RBF at
1000 rows and 32 features from 85.0 ms to 3.0 ms, LinearRegression at 10000
rows and 32 features from 47.8 ms to 2.6 ms. It now comes from run 36166087543,
which ran this branch's library code.

The gate also had an absolute floor of 0.10 ms, which is smaller than what a
shared runner does to a row that size. KernelSVC_RBF predict at 100 rows
measured 0.147, 0.160 and 0.169 ms on three consecutive runs of the same code
and 0.322 ms on the fourth, which failed the build. The floor is now 0.25 ms.
Checked that a 1.5x slowdown on RandomForest at 10000 rows still fails it.

README now cites CI's own measurement of the scaled matrix, 34 rows won of 34
on an Intel Xeon with OpenBLAS, rather than a developer machine's.
@godofecht
godofecht merged commit e965c78 into main Sep 25, 2026
10 checks passed
@godofecht
godofecht deleted the scale-parity branch September 25, 2026 18:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant