Repository navigation
Win every scaled benchmark row - #506
Merged
Merged
Conversation
The 19 canonical rows all fit in 1797 samples. The scaled matrix runs five estimators at 100, 1000 and 10000 rows by 8 and 32 features, and four of its 34 rows were losses, the worst at 0.24x. All four were algorithmic rather than constant-factor. Coordinate descent never terminated early. The stopping test compared the epoch's largest coefficient change against 1e-16, an exact-fixpoint condition that floating point does not reach, so every fit ran its full 1000 epochs where sklearn's cd_fast takes three to ten. Replaced with sklearn's relative criterion, d_w_max / w_max < tol. Lasso at 10000x32 goes 0.24x to 2.50x and diabetes 2.60x to 7.36x. Coefficients move closer to sklearn's. The diabetes parity residual drops from 1.168e-05 to 7.809e-06. LinearRegression solved every design by Householder QR. Tall well-conditioned designs now go through the normal equations with dgemm and dpotrf, guarded by a diagonal-ratio check that falls back to QR when the Gram matrix is near singular. 10000x32 goes 0.76x to 5.35x, diabetes to 10.21x. The RBF SMO built the full n-by-n kernel up front. Rows are now computed on demand through one sgemv and cached. KernelSVC_RBF at 1000x8 goes 0.51x to 1.36x. The Lasso Gram build is routed through dgemm and dgemv above a size product of 65536, with the scalar path kept below it. Scaled: 34 wins, 0 losses, measured on a quiet machine at load 5. Canonical stays 19/19. Suite 114/114, all 11 validators green. PCA on iris is the only other parity row that moved, 6e-08 to 0. tests/test_opt_lasso_fit.flow carried its own copy of the old stopping rule, so it was comparing two different algorithms. Its reference now uses the relative criterion, with a comment saying why.
SMO asks the kernel cache for about a hundred rows per fit, and each row was one cblas_sgemv. At 800 samples and 8 features a row is 6400 flops, which is under the point where a threaded sgemv earns the cost of handing work to its threads. An unrolled scalar dot product below a size product of 65536 takes the whole fit from 0.96 ms to 0.37 ms at 8 features and from 1.04 ms to 0.63 ms at 32, best of seven on an Apple M4 Max at load 5. Above the gate the row is large enough for BLAS to pay and nothing changes: digits at 1437 samples and 64 features is 91968 and stays on the sgemv path. Two other shapes were tried and dropped. Folding the label product into the row build saves a second pass and an allocation per row, worth 9 percent on a binary fit, but it requires each one-vs-one pair to build its own cache, and that costs 3.2x on ten classes at digits scale: 18.0 ms to 58.4 ms. Sharing kernel rows across pairs is worth more than the fused pass. A plain scalar loop without the four accumulators is slower than sgemv at 32 features. Both KernelSVC_RBF rows of the scaled matrix lose on the CI runner, at 0.68x and 0.88x, while winning on this machine. This is the first change aimed at that gap.
A one-vs-one fit kept the whole training matrix and a coefficient matrix with one row per training sample, so every decision value evaluated the RBF kernel against all of it and multiplied most of the result by zero. Most rows carry a zero dual coefficient in every pair. The fit now compacts to the rows some pair gave a nonzero coefficient, which is what libsvm stores. Predict on 200 test rows against 800 training samples goes from 0.58 ms to 0.10 ms, and the whole scaled row from 1.26 ms to 0.83 ms at 8 features and 1.54 ms to 1.18 ms at 32, best of seven on an Apple M4 Max at load 2. Against scikit-learn's 1.37 ms and 1.78 ms on the same machine that is 1.65x and 1.51x, from 0.93x and 1.16x. A fit that finds no support vector keeps one zeroed row, so predict still has a shape BLAS takes and every decision value comes back as its bias.
test_opt_svc_predict asserted that a fitted KernelSVCMulti records the whole training set, which no longer holds now that the fit compacts to support vectors. It now checks the model keeps between one row and the training set size. The assertion that matters is unchanged: the one-vs-one decision values still match the per-pair reference path, to 1.1e-06 on iris.
benchmarks/scaled_flow_baseline.json dated from GitHub Actions run 32002315336 and every row had moved since, most by an order of magnitude: KernelSVC_RBF at 1000 rows and 32 features from 85.0 ms to 3.0 ms, LinearRegression at 10000 rows and 32 features from 47.8 ms to 2.6 ms. It now comes from run 36166087543, which ran this branch's library code. The gate also had an absolute floor of 0.10 ms, which is smaller than what a shared runner does to a row that size. KernelSVC_RBF predict at 100 rows measured 0.147, 0.160 and 0.169 ms on three consecutive runs of the same code and 0.322 ms on the fourth, which failed the build. The floor is now 0.25 ms. Checked that a 1.5x slowdown on RandomForest at 10000 rows still fails it. README now cites CI's own measurement of the scaled matrix, 34 rows won of 34 on an Intel Xeon with OpenBLAS, rather than a developer machine's.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The 19 canonical rows all fit in 1797 samples. The scaled matrix runs five estimators at 100, 1000 and 10000 rows by 8 and 32 features, and four of its 34 rows were losses, the worst at 0.24x. All four were algorithmic.
Coordinate descent never terminated early. The stopping test compared the epoch's largest coefficient change against
1e-16, an exact-fixpoint condition that floating point does not reach, so every fit ran its full 1000 epochs where sklearn'scd_fasttakes three to ten. Replaced with sklearn's relative criterion,d_w_max / w_max < tol. Lasso at 10000x32 goes 0.24x to 2.50x and diabetes 2.60x to 7.36x. Coefficients move closer to sklearn's. The diabetes parity residual drops from 1.168e-05 to 7.809e-06.LinearRegression solved every design by Householder QR. Tall well-conditioned designs now go through the normal equations with
dgemmanddpotrf, guarded by a diagonal-ratio check that falls back to QR when the Gram matrix is near singular. 10000x32 goes 0.76x to 5.35x, diabetes to 10.21x.The RBF SMO built the full n-by-n kernel up front. Rows are now computed on demand through one
sgemvand cached. KernelSVC_RBF at 1000x8 goes 0.51x to 1.36x.The Lasso Gram build is routed through
dgemmanddgemvabove a size product of 65536, with the scalar path kept below it.Results
Scaled numbers measured locally on an Apple M4 Max at load average 5, three repeats aggregated by median, with the sklearn baseline regenerated in the same session. CI regenerates the same matrix and reports it without gating.
PCA/irisis the only other parity row that moved, 6e-08 to 0. The other 17 are unchanged.tests/test_opt_lasso_fit.flowcarried its own copy of the old stopping rule, so it was comparing two different algorithms. Its reference now uses the relative criterion, with a comment saying why.benchmarks/scaled_flow_baseline.jsonis left alone. It is a Flow-only self-regression gate from a GitHub Actions artifact, it is now 1.16x to 168x behind the current code, and refreshing it from this laptop would make a fast machine the standard CI has to meet. Refresh it from a CI artifact of this run instead. README says so.