Preprocessing encoding - CyrilB1531/lodestar GitHub Wiki

Development build. This page describes main, not a released package. The latest published Lodestar.Preprocessing is 0.1.0 — read its documentation.

HomePreprocessing

Encoding and imputation — Lodestar.Preprocessing

Two encoders and an imputer, at sklearn.preprocessing and sklearn.impute parity, over row-major spans rather than an IDataView: Encoders.OneHot gives each category a column, Encoders.Ordinal gives it a code, and SimpleImputer fills what is missing.

One element type per call. A 2-D array carries one dtype in the reference too, so the encoders are generic over the category type and a caller with a string column and an integer column makes two calls. The type decides the order — strings sort by code point as numpy's do, integers as numbers — and that order decides the columns.

Why this exists when ML.NET has all three

ML.NET has OneHotEncoding, MapValueToKey and ReplaceMissingValues; SharpLearning has OneHotTransformer and ReplaceMissingValuesTransformer. Nothing here is absent from .NET. Each one of them is reached through an IDataView or through a catalog naming columns, or works on that library's own matrix type — and what is absent is a call that takes an array and returns one. That is the whole argument, and decision 0004 says so rather than claiming a capability gap.

Types

Type What it is
Encoders Fits the two encoders.
OneHotEncoder One column per category.
OneHotEncoderOptions Which category to drop, and what an unseen value becomes.
CategoryDrop None, the first, or the first of a binary feature.
UnknownCategory Refuse an unseen value, or encode it as zeros.
OrdinalEncoder One code per category.
SimpleImputer Fills missing values with a per-feature statistic.
SimpleImputerOptions Which statistic, and what to do with an empty feature.
ImputationStrategy Mean, median, most frequent, or a constant.

See also