Skip to content
Jennifer Programming Language

ml - classical machine learning

Enable with use ml;. Classical / predictive machine learning on tabular data - the scikit-learn-lite core companion to stats and linalg. Models follow a fit / predict shape: a fit function (ml.kMeans, ml.linearRegression, ...) trains and returns an opaque ml.Model handle, and ml.predict / ml.transform apply it. Pure Go stdlib, TinyGo-clean, both binaries.

It is not a deep-learning framework: tensors, autodiff, and training deep nets are native GPU / C++ territory a tree-walker cannot usefully replace, and "run a pre-trained model" is already http / os.run.

jennifer
use ml;
use io;

# Fit a model, then apply it.
def X as list of list of float init [[1.0, 1.0], [2.0, 1.0], [1.0, 2.0], [3.0, 2.0]];
def y as list of float init [6.0, 8.0, 9.0, 13.0];   # y = 2*x1 + 3*x2 + 1
def model as ml.Model init ml.linearRegression($X, $y);
io.printf("%v\n", ml.predict($model, [[5.0, 5.0]]));  # ~[26.0]

Data shape

A feature matrix X is a list of list of float/int (one inner list per sample, one entry per feature); labels / targets y are a list of float/int. Class labels are integers (0, 1, 2, ...); regression targets are any number. ml.predict returns a list of int for a classifier / cluster model and a list of float for a regressor.

Fitting models

Each returns an opaque ml.Model handle. A fitted model is immutable, so a handle is safe to share across value-copies and spawned tasks (read-only).

Fit callKindNotes
ml.linearRegression(X, y)regressionOrdinary least squares (normal equations).
ml.ridge(X, y, alpha)regressionL2-regularized OLS; alpha >= 0 shrinks the coefficients (not the intercept).
ml.lasso(X, y, alpha)regressionL1-regularized (coordinate descent); drives small coefficients to exactly 0.
ml.kNN(X, y, k)classifierk-nearest-neighbours majority vote (Euclidean).
ml.kNNRegressor(X, y, k)regressionk-NN averaging the k nearest targets.
ml.naiveBayes(X, y)classifierGaussian naive Bayes (multiclass).
ml.logisticRegression(X, y [, lr [, epochs]])classifierBinary or multiclass (one-vs-rest); gradient descent (lr default 0.1, epochs 1000). Only the binary case has predictProba.
ml.decisionTree(X, y [, maxDepth])classifierCART with Gini impurity (maxDepth default 8).
ml.decisionTreeRegressor(X, y [, maxDepth])regressionCART with variance-reduction splits, leaf = mean target.
ml.randomForest(X, y [, nTrees [, maxDepth]])classifierBagged trees with per-split feature subsampling (nTrees 10).
ml.randomForestRegressor(X, y [, nTrees [, maxDepth]])regressionBagged regression trees (mean of tree predictions).
ml.kMeans(X, k [, maxIter])clusteringLloyd's algorithm, k-means++ seeding.
ml.pca(X, nComponents)transformPrincipal component analysis (covariance eigendecomposition).
ml.standardScaler(X)transformFit a z-score scaler (per-feature mean / stddev).
ml.minMaxScaler(X)transformFit a [0, 1] min-max scaler.

Inspecting a fitted model

A model is opaque, but you can read the parameters it learned:

CallApplies toReturns
ml.coefficients(model)linear / ridge / lasso / logisticlist of float (per feature), or list of list of float (per class) for multiclass logistic.
ml.intercept(model)linear / ridge / lasso / logisticfloat, or list of float per class for multiclass logistic.
ml.centroids(model)kMeanslist of list of float (cluster centres).
ml.components(model)pcalist of list of float (principal axes, one row per component).
ml.explainedVariance(model)pcalist of float - the variance ratio each component captures (choose nComponents from these).
ml.featureImportances(model)decisionTree / randomForestlist of float (Gini importances, summing to 1).
jennifer
use ml;
use io;
def m as ml.Model init ml.linearRegression([[1.0, 1.0], [2.0, 1.0], [1.0, 2.0], [3.0, 2.0]], [6.0, 8.0, 9.0, 13.0]);
io.printf("y = %v . x + %v\n", ml.coefficients($m), ml.intercept($m));   # [2, 3], 1

Applying a model

CallReturns
ml.predict(model, X)list of int / list of floatPredicted labels (classifier / cluster) or values (regressor) for each row of X.
ml.transform(model, X)list of list of floatTransformed features (scalers, PCA).
ml.predictProba(model, X)list of floatPositive-class probability, for logisticRegression.
ml.free(model)nullDrop the model to free its memory early (a handle otherwise lives for the run).

The random models (kMeans, randomForest, trainTestSplit) draw from math's shared random source, so math.randSeed(n) makes a run reproducible. This is deliberate - reproducibility, not secrecy, is what training needs; a seed is not a secret, so ml uses math's seedable source, never crypto's unseedable one (which would make a split or clustering impossible to reproduce).

A fitted model lives in a per-run registry behind its handle until the program ends. When you fit many models in a loop (e.g. a large cross-validation), call ml.free(model) on the ones you are done with so the registry does not grow unbounded. The cost-driving hyper-parameters are bounded (tree maxDepth <= 64, forest nTrees <= 1000, kMeans maxIter <= 10000, logistic epochs <= 1e6, kFold nSamples * k <= 1e8, multiclass logistic <= 100 classes, polynomialFeatures degree <= 8, output width <= 1e5, and total cells (rows x columns) <= 2e7); a value above the ceiling is a catchable error, as is a polynomialFeatures product that overflows to a non-finite value.

ml.lasso, unlike ordinary least squares, is fine with more features than rows (p > n) - that is its feature-selection use case, so it does not require rows > features. Feeding it a very wide polynomialFeatures design is allowed but does dense O(sweeps x features x rows) work; keep the expansion modest. ml.logLoss requires each probability in [0, 1] (an out-of-range value errors, rather than being silently clamped).

jennifer
use ml;
use math;
use io;

math.randSeed(1);
def data as list of list of float init [[0.0, 0.0], [0.2, 0.1], [5.0, 5.0], [5.1, 4.8]];
def km as ml.Model init ml.kMeans($data, 2);
io.printf("clusters: %v\n", ml.predict($km, $data));   # e.g. [0, 0, 1, 1]

# Scale features, then reduce dimensionality.
def scaler as ml.Model init ml.standardScaler($data);
def scaled as list of list of float init ml.transform($scaler, $data);

Model selection

CallReturns
ml.trainTestSplit(X, y, testFraction)ml.SplitShuffle and split into {trainX, trainY, testX, testY}; testFraction in (0, 1).
ml.kFold(nSamples, k)list of ml.Foldk contiguous folds, each {trainIdx, testIdx} (index lists into your data).
ml.polynomialFeatures(X, degree)list of list of floatExpand to all monomials up to degree (with a leading bias column); stateless, so the same call fits train and test.
jennifer
export def struct Split { trainX as list of list of float, trainY as list of float, testX as list of list of float, testY as list of float };
export def struct Fold  { trainIdx as list of int, testIdx as list of int };

Metrics

Pure functions over label / value lists (no model). Classification metrics take (yTrue, yPred); precision / recall / f1 take an optional third positive-label argument (default 1).

CallMeaning
ml.accuracy(yTrue, yPred)Fraction of exact matches.
ml.precision(yTrue, yPred [, positive])TP / (TP + FP) for the positive label.
ml.recall(yTrue, yPred [, positive])TP / (TP + FN).
ml.f1(yTrue, yPred [, positive])Harmonic mean of precision and recall.
ml.confusionMatrix(yTrue, yPred)ml.Confusion{labels, matrix} (rows = true, columns = predicted).
ml.rocAuc(yTrue, scores)Binary ROC-AUC from 0 / 1 labels and predicted scores (tie-aware).
ml.logLoss(yTrue, probas)Binary cross-entropy of 0 / 1 labels against predicted probabilities.
ml.rmse(yTrue, yPred) / ml.mse / ml.maeRoot-mean-square / mean-square / mean-absolute error (regression).
ml.r2(yTrue, yPred)Coefficient of determination R^2 (zero-variance targets error).
jennifer
use ml;
use io;
def yt as list of float init [1.0, 0.0, 1.0, 1.0, 0.0, 1.0];
def yp as list of float init [1.0, 0.0, 0.0, 1.0, 0.0, 1.0];
io.printf("accuracy %v, F1 %v\n", ml.accuracy($yt, $yp), ml.f1($yt, $yp, 1));

Strictness and scope

Like the rest of the numeric stack, a degenerate input is a catchable error, not a silent NaN: an empty / ragged matrix, mismatched X / y lengths, a singular regression design (collinear features), a diverged logistic fit, or magnitudes that overflow the computation all raise a positioned error.

ml targets the modest tabular data a tree-walker handles - native loops over thousands of rows, not millions. Large-scale training stays a native-tool job. Support-vector machines, gradient-boosted trees, density-based clustering (DBSCAN), and model serialization are out of scope for this tier.

See also

stats (distributions + inference - the estimate / test companion to ml's fit / predict), linalg (linear algebra), and math. The module index.