Stats 0.1.0 hypothesis testing - CyrilB1531/lodestar GitHub Wiki
Lodestar.Stats 0.1.0. This page is frozen at that release. Read the current documentation for what
mainsays now. A link to a decision or a migration page followsmain, and leaves the archive.
Hypothesis testing
Lodestar.Stats answers one question in ten forms: is this difference more
than noise?
Which test
| you have | and you assume | use |
|---|---|---|
| two independent samples | roughly normal | TTest.Independent |
| two independent samples | nothing about the shape | MannWhitney.Test |
| the same subjects measured twice | roughly normal differences | TTest.Paired |
| the same subjects measured twice | nothing about the shape | Wilcoxon.Paired |
| counts in categories | a stated expected distribution | ChiSquare.GoodnessOfFit |
| a contingency table | cells large enough for the approximation | ChiSquare.Contingency |
| a 2ร2 table with small cells | nothing | FisherExact.Test |
| two samples, whole distributions | nothing | KolmogorovSmirnov.TwoSample |
| three or more groups | roughly normal, similar spread | OneWayAnova.Test |
| three or more groups | nothing about the shape | KruskalWallis.Test |
| one sample, and a normality assumption to check | nothing | ShapiroWilk.Test |
| many p-values at once | nothing | MultipleComparisons |
What a p-value is, and is not
A p-value is the probability of seeing a difference at least this large if the null hypothesis is true. It is not the probability that the null hypothesis is true, and it is not the probability that your result is a fluke. A p-value of 0.03 does not mean there is a 3 % chance you are wrong.
Two consequences worth acting on:
0.049and0.051are the same evidence. The threshold is a convention, not a discovery. Report the number.- Twenty tests at 5 % produce one significant result by chance. That is what
MultipleComparisonsis for, and it is not optional once you are testing more than a couple of things.
The step-up procedure below is
MultipleComparisons.BenjaminiHochberg,
the one most people reach for first:
using Lodestar.Stats;
double[] pValues = [0.001, 0.008, 0.039, 0.041, 0.042];
double[] adjusted = MultipleComparisons.BenjaminiHochberg(pValues);
bool stillSignificant = adjusted[0] < 0.05; // => True
One default that is not scipy's
TTest.Independent defaults to
Welch's test; scipy.stats.ttest_ind defaults to Student's. Pooling the two
variances is only correct when the populations really share one, which is an
assumption most callers have not checked. Pass Variance.Equal for scipy's
default. Everything else in this package matches scipy.stats 1.18.0 exactly,
and the equivalence table is the row-by-row map, including
the handful of places one call refuses rather than answering โ a NaN, a
warning, or a number a caller did not ask for.
Exact and asymptotic
Three tests carry both an exact null distribution and a normal approximation to
it, selected by ExactMethod:
Autoโ exact for a small, untied sample; asymptotic otherwise. Whatscipy'smethod='auto'does, and the same thresholds.Exactโ always the exact distribution. On tied data the number is only approximate, because ties break the equal-probability argument the enumeration rests on;scipycomputes there too rather than refusing, and so does this.Asymptoticโ always the normal approximation, whatever the sample size.
The branch changes the number, not just the running time, which is why it is a
parameter and not a hidden optimisation. Each of the three tests also refuses
ExactMethod.Exact past its own size bound โ the exact table is
O(nยทm) or worse to build, and Auto never crosses that bound on its own,
falling back to the asymptotic answer instead. The three reference pages
(MannWhitney.Test,
Wilcoxon.Paired,
KolmogorovSmirnov.TwoSample)
each state their own bound, because it is not the same number twice.
No maintained incumbent to compare against
MathNet.Numerics 5.0.0 (2022-04-03, 74.7M downloads) is the dominant
third-party numerical library for .NET, and it ships probability distributions
and descriptive statistics โ no hypothesis tests. Accord.Statistics 3.8.0,
the one .NET library that did carry them, was last published on
2017-10-19; its framework, accord-net/framework (4.5k stars), was
archived by its owner on 2020-11-19. ML.NET does prediction: a t-test
exists only in Azure ML Studio (classic), a retired hosted product, and
Mann-Whitney only in Kusto/KQL โ neither is a .NET library a project can
reference. There is nothing maintained to benchmark against, which is itself
the finding, the same shape as Lodestar.Conformal's survey โ
but unlike Lodestar.Conformal's case, Accord.Statistics is still
installable, and #442's own constraint asks for a named .NET incumbent where
one exists at all. TTest.Independent,
MannWhitney.Test and
ChiSquare.Contingency are benchmarked and
cross-checked against it in
bench/README.md
and docs/guides/performance.md โ
archived is not the same as absent, and this package's own oracle discipline
means the comparison is also a second opinion on scipy, not only a timing.