Keyboard shortcuts

Press or to navigate between chapters

Press ? to show this help

Press Esc to hide this help

Profiling Guide

How to profile numeric, categorical, and mixed data, and how to act on the data-quality findings. Types and functions live in datarust_profile; the statistics they report flow from datarust::stats.

All public APIs return Result; invalid input is a recoverable ProfileError, never a panic.

Profiling numeric data

Pass a Matrix to profile_matrix. Non-finite values (NaN, ±inf) are treated as missing and counted separately — they never poison the mean or standard deviation.

#![allow(unused)]
fn main() {
use datarust::Matrix;
use datarust_profile::profile_matrix;

let m = Matrix::from_rows(vec![
    vec![1.0, 100.0],
    vec![2.0, 110.0],
    vec![3.0, 105.0],
    vec![4.0, 95.0],
    vec![5.0, f64::NAN], // missing reading in column 1
])?;

let p = profile_matrix(&m, Some(&["index".into(), "reading".into()]))?;
}

Each numeric column yields a NumericStats:

FieldMeaning
mean, stdCentral tendency and spread (sample std, ddof = 1).
fiveFiveNumber: min / Q1 / median / Q3 / max.
skewnessFisher–Pearson third moment. 0 ≈ symmetric; positive ⇒ right tail.
kurtosisExcess (Fisher) fourth moment. 0 ≈ normal; positive ⇒ heavy tails.
histogramHistogram with Sturges-bin edges and counts.
outlier_count, outlier_fractionValues beyond the Tukey IQR fences (Q1 − 1.5·IQR, Q3 + 1.5·IQR).

The histogram bin count follows Sturges’ rule (ceil(log2(n) + 1), floored at 1). Read histogram.counts directly, or use histogram.nbins() / histogram.max_count() for rendering.

Fast path

profile_matrix routes the mean, variance, and five-number summary through datarust’s fused flat-buffer helpers (column_mean_var_flat, column_quantiles_many_flat) over Matrix::as_slice(), avoiding per-column Vec allocation on wide tables. Columns containing NaN fall back to the per-column path, since the flat helpers are not NaN-aware.

Profiling categorical data

profile_str_matrix infers each column’s type. A column is Numeric when every non-empty cell parses as f64; otherwise it is Categorical. Empty markers recognised as missing: "", NA, N/A, null, NaN, None, -, ?.

#![allow(unused)]
fn main() {
use datarust::StrMatrix;
use datarust_profile::profile_str_matrix;

let s = StrMatrix::from_strings(vec![
    vec!["Istanbul", "basic"],
    vec!["Ankara", "premium"],
    vec!["Izmir", "basic"],
    vec!["Istanbul", "basic"],
])?;

let p = profile_str_matrix(&s, Some(&["city".into(), "tier".into()]))?;
}

A CategoricalStats reports unique (cardinality), top / freq (most frequent value and its count), imbalance_ratio (freq ÷ non-missing cells), and top_values — the top-N (value, count) pairs, sorted descending, for frequency charts.

Mixed tables

Real datasets mix numeric and categorical columns. profile_table takes an optional numeric [Matrix] and an optional categorical [StrMatrix] side by side, with one shared row count. names must list every column across both blocks (numerics first, then categoricals).

#![allow(unused)]
fn main() {
use datarust_profile::profile_table;

let p = profile_table(
    Some(&numeric),
    Some(&categorical),
    &["age".into(), "income".into(), "city".into(), "tier".into()],
)?;
}

Duplicate-row detection works across both blocks: a row is a duplicate only if its numeric and categorical cells match an earlier row exactly.

Pairwise relationships & target leakage

v0.3 computes pairwise interactions across columns:

  • Pearson correlation matrix (rels.pearson): computed over numeric columns using datarust::stats::correlation_matrix.
  • Cramér’s V matrix (rels.cramers_v): computes categorical association (V ∈ [0.0, 1.0]) via a pure-Rust contingency table.
  • Point-biserial correlation (rels.point_biserial): measures correlation between binary categorical columns and continuous numeric columns.

Target-leakage detection

To check for feature leakage against a target column, construct the profile with profile_matrix_with_target or profile_table_with_target:

#![allow(unused)]
fn main() {
use datarust_profile::profile_table_with_target;

let p = profile_table_with_target(
    Some(&numeric),
    Some(&categorical),
    &["age".into(), "income".into(), "city".into(), "churn".into()],
    "churn",
)?;
}

If a feature column has high correlation or association (|r| >= 0.90 or V >= 0.90) with the target column, run_checks fires a QualityKind::TargetLeakage issue.

Data quality checks

run_checks turns a profile into a list of QualityIssues. The thresholds are conservative defaults; tune them via Thresholds.

#![allow(unused)]
fn main() {
use datarust_profile::quality::{Thresholds, QualityKind};
use datarust_profile::quality::checks::run_checks;

let mut t = Thresholds::default();
t.outlier_fraction = 0.02;  // flag columns with ≥2% outliers
t.high_correlation = 0.90;  // flag numeric pairs with |r| >= 0.90
t.target_leakage = 0.85;    // flag feature-target correlation >= 0.85

for issue in run_checks(&p, &t) {
    println!("{:?} [{}] {}: {}",
        issue.kind, issue.severity,
        issue.column.as_deref().unwrap_or("(dataset)"),
        issue.message);
}
}

The eight checks

KindScopeFires when
HighMissingcolumnmissing_fraction ≥ threshold.missing_fraction (default 0.5). Severity escalates to Critical at 0.9.
ConstantColumncolumn (numeric)variance ≤ threshold.near_zero_variance (default 1e-12).
NearUniquecolumn (categorical)unique / count ≥ threshold.near_unique_ratio (default 0.98) — likely an identifier, not a feature.
Outlierscolumn (numeric)outlier_fraction ≥ threshold.outlier_fraction (default 0.05). Severity escalates to Warning at 0.2.
Imbalancecolumn (categorical)imbalance_ratio ≥ threshold.imbalance_ratio (default 0.95). Always Critical.
HighCorrelationcolumn pair (numeric)Pearson |r| ≥ threshold.high_correlation (default 0.95). Collinearity risk.
TargetLeakagecolumn (feature)Feature-target correlation or Cramér’s V ≥ threshold.target_leakage (default 0.90). Always Critical.
DuplicateRowsdatasetduplicate_rows > 0. Severity escalates to Warning at 0.1.

Each QualityIssue carries a Severity (Info, Warning, Critical) and an optional column name (dataset-wide findings like DuplicateRows set column to None).

Rendering reports

The report module renders a profile plus its findings. HTML is always available; JSON needs the serde feature.

#![allow(unused)]
fn main() {
use datarust_profile::report;

// HTML: self-contained, no dependencies.
let html = report::to_html(&p);

// Bring your own findings (e.g. with custom thresholds):
let findings = run_checks(&p, &t);
let html = report::to_html_with(&p, &findings);

// JSON (serde feature):
#[cfg(feature = "serde")]
let json = report::to_json(&report::JsonReport::from_profile(&p))?;
}

The HTML report uses a responsive card grid: numeric cards carry summary statistics, CSS mini-histograms, and outlier counts; categorical cards carry top values, imbalance ratios, and frequency bar charts. A Relationships section presents correlation heatmaps for numeric (Pearson) and categorical (Cramér’s V) pairs, plus point-biserial tables. Findings appear in a severity-coloured list at the top.