People Analytics Toolkit: 50 Production-Grade Feature Engineering Formulations

Published October 2026 • Author: Sam Tritto • Package: people-analytics-toolkit (v0.1.0)
Software Engineering Machine Learning Statistics Bayesian
← Back to More Projects
View on PyPI (v0.1.0) View on GitHub
People Analytics Toolkit Logo
people-analytics-toolkit: A production-grade Python library for workforce intelligence, causal uplift, and talent feature engineering.

Introduction: Why People Analytics Needs Dedicated Feature Engineering

Standard machine learning tutorials often assume independent and identically distributed (i.i.d.) tabular observations. Real-world human capital management (HCM) systems violate these assumptions at every scale. Associate records exist inside nested organizational hierarchies, experience non-linear psychological decay, suffer from unobserved managerial frictions, and trigger viral contagion effects across collaborative networks.

When standard data science workflows feed raw tenure, department tags, and salary bands into generic gradient-boosted trees, the models routinely fail in production. They overfit small cohorts, suffer target leakage from high-cardinality department codes, miss temporal tipping points, and reinforce historical pay disparities.

I engineered people-analytics-toolkit to bridge this divide. It provides an enterprise-ready, mathematically rigorous library containing 50 specialized feature engineering formulations structured across 11 operational pillars. This document serves as the comprehensive technical reference for the package following its official v0.1.0 release on PyPI.

Data Transparency & Privacy: All empirical visualizations, benchmark telemetry, and code executions demonstrated in this guide utilize synthetically generated enterprise cohorts. The data generators replicate realistic retail, healthcare, and technology human capital properties (such as power-law store hierarchies, tenure hazard curves, and collaboration networks) without exposing proprietary records or personally identifiable information (PII).
HR Analytics in Practice - Toby Flenderson
Why is HR the way that it is? Elevating Human Resources from administrative burden into predictive, data-driven workforce intelligence.

Installation and Core Dependencies

The people-analytics-toolkit package (v0.1.0) is built for modern Python (3.11+) and integrates directly with the standard scientific Python ecosystem. You can install it directly from PyPI:

From PyPI

Install the core framework or include optional feature engineering bundles in square brackets:


pip install people-analytics-toolkit

# All optional feature engineering bundles
pip install "people-analytics-toolkit[all]"

# Or select specific domain bundles
pip install "people-analytics-toolkit[timeseries,explainability]"

Optional Dependency Bundles

To keep the base installation lightweight and fast, specialized mathematical and domain-specific dependencies are organized into modular extras:

Bundle Extra Included Libraries Features & Capabilities Enabled
[timeseries] arch, hmmlearn GARCH(1,1) dynamic volatility indices, Hidden Markov Model (HMM) latent regimes
[anomalies] pyod Local Outlier Factor (LOF) and density-based anomaly detectors
[deeplearning] torch, pyod Autoencoder reconstruction error, deep trajectory embeddings
[survival] lifelines Kaplan-Meier & Nelson-Aalen hazard rate embeddings, survival curves
[explainability] shap, lightgbm 2D SHAP interaction matrices, turnover risk inflection thresholds
[all] All of the above Complete 50-feature analytical and modeling suite
[dev] pytest, hypothesis, jupyterlab, mypy, matplotlib, seaborn Development, property testing, CI validation, and interactive notebooks

Core Foundation Libraries

The base package installation without extras provides zero-overhead access to core mathematical, econometric, and organizational network formulations:

  • pandas & numpy: Vectorized matrix math, temporal windowing, and high-performance cohort group operations.
  • scipy & statsmodels: Econometric estimation, empirical Bayes shrinkage, SVD decomposition, and time-series kernel smoothing.
  • scikit-learn: Estimators, cross-validation splits, and manifold distance metrics.
  • networkx: Organizational network graph traversal, centrality decomposition, and random walk generation.

Local Development with uv

If you are contributing to the codebase, running benchmarks, or validating CI tests locally, synchronize dependencies using uv:

cd people-analytics-toolkit

# Sync with all 50-feature mathematical dependencies:
uv sync --extra all

# Or sync all dependencies including testing, JupyterLab, and dev tooling:
uv sync --all-extras

Complete Architectural Taxonomy: 50 Features Across 11 Pillars

The 50 features in people-analytics-toolkit are categorized into 11 distinct problem pillars addressing specific organizational phenomena:

Pillar Operational Domain Features Included
1. High Cardinality & Hierarchies Cohort size disparity, sparse departments, leakage
2. Preserving Memory & Lifecycles Long-range temporal decay, non-stationary headcounts
3. Chaos, Volatility & Regimes Scheduling turbulence, variance clustering, contagion
4. Structural Anomalies & Sequences Outlier profiles, career sequence divergence, drift
5. Attribution, Interactions & Accounting Rate-mix shifts, synergistic risks, hypothesis audits
6. Career Mobility & Tipping Points Graph transitions, stagnation spikes, flight risk surges
7. Compensation & Pay Equity Band parity, salary compression, red/green circling
8. Relational Dynamics & ONA Information bottlenecks, structural holes, manager load
9. Algorithmic Equity & Fairness Pay gap decomposition, equal opportunity, Lipschitz
10. Prescriptive Causal Interventions Confounded observational policies, uplift targeting
11. Continuous Trajectories & Graph Learning Holding durations, multi-relational skills, attention

Pillar 1: High Cardinality & Hierarchical Cohorts

In HCM data, associates are grouped into hierarchical clusters: stores, teams, departments, job codes, and cost centers. Small cohorts have high variance; naive mean encoding creates fatal out-of-sample target leakage. Pillar 1 stabilizes small-sample cohorts through empirical Bayes shrinkage and leave-one-out adjustments.

Feature 1: Bayesian Target Encoding (Empirical Bayes Shrinkage)

Business Context: Small retail stores or specialized functional teams often exhibit extreme historical turnover rates (0% or 100%) purely due to small sample size. Using raw cohort averages leads to catastrophic overfitting. Bayesian target encoding shrinks small cohort averages toward the enterprise prior mean based on a credibility weight \(m\).

Mathematical Formulation:

$$\hat{S}_i = \lambda(n_i) \cdot \bar{y}_i + (1 - \lambda(n_i)) \cdot \bar{y}_{\text{global}}, \quad \lambda(n_i) = \frac{n_i}{n_i + m}$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.cardinality import BayesianTargetEncoder

# Instantiate encoder with smoothing prior and cross-validation folds
bte = BayesianTargetEncoder(
    m=15.0,                 # Prior smoothing weight / pseudocount (e.g., 5.0, 10.0, 15.0, 25.0, 50.0)
    cv_folds=5,             # K-fold splits for out-of-fold leakage prevention (e.g., 3, 5, 10, None for in-sample)
    seed=42                 # Random seed for fold shuffling
)

# Compute out-of-fold target encoding for enterprise training rosters
encoded_series = bte.fit_transform_oof(
    df=roster_df,           # pandas DataFrame containing cohort and target columns
    group_col='store_dept', # Column name or composite hierarchy key (e.g., 'store_id', 'department_id', 'store_dept')
    target_col='turnover'   # Target binary indicator (0/1) or continuous outcome column
)

# In-sample transform for out-of-sample validation or test sets
test_encoded = bte.transform(
    df=val_roster_df,
    group_col='store_dept'
)
Bayesian Target Encoding shrinkage curve
Figure 1: Empirical Bayes shrinkage curves. Small cohorts (\(n < 15\)) are regularized heavily toward the enterprise base rate, while large cohorts retain their empirical group mean.

Feature 2: Grouped Leave-One-Out (LOO) Z-Scores

Business Context: To assess whether an employee is logging excessive overtime or underperforming relative to immediate peers, one must calculate a peer z-score. However, if the employee's own value is included in the cohort mean and standard deviation, the metric is biased toward zero in small teams. LOO Z-scores remove the focal individual from the peer cohort statistics.

Mathematical Formulation:

$$\mu_{-i} = \frac{\sum_{j \neq i} x_j}{n - 1}, \quad \sigma_{-i} = \sqrt{\frac{\sum_{j \neq i} (x_j - \mu_{-i})^2}{n - 2}}, \quad z_i = \frac{x_i - \mu_{-i}}{\sigma_{-i} + \epsilon}$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.cardinality import GroupedLOOZScore, calculate_grouped_loo_zscore

# Object-oriented pipeline approach
loo_calculator = GroupedLOOZScore(
    group_col='department_id',   # Cohort grouping key (e.g., 'team_id', 'department_id', 'store_id')
    target_col='overtime_hours', # Continuous telemetry metric (e.g., 'overtime_hours', 'sales_quota', 'absence_days')
    ddof=1,                     # Delta degrees of freedom (0 for population variance, 1 for sample variance)
    eps=1e-8                    # Epsilon floor to prevent division-by-zero on zero-variance cohorts
)
loo_z_series = loo_calculator.transform(df=timeseries_df)

# Or functional execution:
loo_z_series = calculate_grouped_loo_zscore(
    df=timeseries_df,
    group_col='department_id',
    target_col='overtime_hours',
    ddof=1,
    eps=1e-8
)
Grouped Leave-One-Out Z-Scores distribution
Figure 2: Distribution of Grouped LOO Z-Scores across operational cohorts, exposing true departmental outliers without focal self-attenuation.

Feature 3: Leave-One-Out Expected Differential (LOO-ED)

Business Context: Evaluating supervisory quality or "manager alpha" requires separating the manager's true impact from the baseline capabilities of their assigned associates. LOO-ED computes the expected performance differential of a cohort when omitting each member, quantifying whether a leader reliably lifts or drags down performance.

Mathematical Formulation:

$$\Delta_{\text{LOO-ED}, i} = y_i - \mathbb{E}[y \mid \text{Cohort}_{-i}]$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.attribution import LeaveOneOutExpectedDifferential

# Quantify manager value-add / cohort differential
loo_ed_model = LeaveOneOutExpectedDifferential(
    group_col='manager_id',      # Cohort or supervisor identifier (e.g., 'manager_id', 'store_id', 'pod_lead')
    metric_col='retention_rate', # Observed performance / retention metric (e.g., 'retention_rate', 'productivity')
    target_col='turnover_rate'   # Benchmark or counterfactual baseline column (e.g., 'turnover_rate', 'baseline_rate')
)

ed_series = loo_ed_model.transform(
    df=shift_manager_performance_df # DataFrame containing shift and cohort observations
)
Leave-One-Out Expected Differential distribution
Figure 3: Shift manager value-add quantified via LOO-ED, isolating supervisor capability from baseline crew performance.

Pillar 2: Preserving Memory, Lifecycles & Time Decay

Workforce behavior is deeply temporal. Events from two weeks ago have vastly different psychological impact than events from twelve months ago. Standard integer differencing (\(d=1\)) wipes out all multi-year organizational memory, while raw metrics suffer from non-stationarity. Pillar 2 implements memory-preserving filters and hazard embeddings.

Feature 4: Fractional Differencing

Business Context: Macro workforce headcount series are typically non-stationary. Taking standard first differences (\(d=1\)) achieves stationarity but destroys all long-memory signals (e.g., seasonal hiring buildup and multi-year restructuring). Fractional differencing expands the lag operator with real order \(d \in (0, 1)\) using binomial series, retaining maximum memory while satisfying stationarity tests (ADF \(p < 0.05\)).

Mathematical Formulation:

$$(1 - L)^d = \sum_{k=0}^{\infty} (-1)^k \binom{d}{k} L^k = 1 - dL + \frac{d(d-1)}{2!} L^2 - \frac{d(d-1)(d-2)}{3!} L^3 + \dots$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.memory import fractional_difference, scan_fractional_differencing

# Compute fractional differencing on workforce headcount or demand series
fd_series = fractional_difference(
    series=headcount_timeseries, # pandas Series, numpy array, or list of longitudinal measurements
    d=0.42,                      # Fractional differencing order (e.g., 0.25, 0.35, 0.42, 0.60)
    threshold=1e-4               # Weight truncation threshold to bound computational memory (e.g., 1e-4, 1e-5)
)

# Automated scan across d values to find minimal d satisfying Augmented Dickey-Fuller stationarity
scan_results_df = scan_fractional_differencing(
    series=headcount_timeseries,
    d_values=None,               # Optional array of d values (defaults to np.linspace(0.0, 1.0, 21))
    threshold=1e-4
)
Fractional Differencing ADF stationarity scan
Figure 4: Fractional differencing memory scan. At \(d=0.42\), the headcount series passes the ADF 95% stationarity threshold while preserving over 78% of original autocorrelation memory.

Feature 5: Weighted Weeks Since (Continuous Event-Decay Features)

Business Context: Employee recognition awards, spot bonuses, and performance citations boost engagement, but their retention effects decay over time. Furthermore, an Executive Excellence Award carries higher lasting value than a peer-to-peer Bronze citation. This feature models continuous exponential decay modulated by award tier multipliers.

Mathematical Formulation:

$$W_i(t) = \sum_{e \in E_i} w_e \cdot \exp\left(-\frac{\ln(2)}{\tau} (t - t_e)\right)$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.memory import weighted_weeks_since

# Calculate continuous decaying recognition capital
decay_df = weighted_weeks_since(
    events_df=rewards_df,          # Event log DataFrame
    current_week=52,               # Current evaluation week / time step (e.g., 52, 104, 156)
    half_life_weeks=8.0,           # Half-life tau in weeks for exponential decay (e.g., 4.0, 8.0, 12.0)
    employee_id_col='employee_id', # Associate identifier column
    event_week_col='award_week',   # Event occurrence week index column
    tier_col='award_tier',         # Categorical award level column
    custom_tier_weights={          # Multiplier weights for award prestige
        'Bronze': 1.0,
        'Silver': 2.5,
        'Gold': 5.0,
        'Executive': 10.0
    }
)
Weighted Weeks Since Event Decay Trajectories
Figure 5: Tier-weighted recognition decay profiles over 52 operational weeks, tracking the psychological half-life of incentive programs.

Feature 6: Exponentially Weighted Moving Average (EWMA) & Sum (EWMS)

Business Context: Simple rolling windows produce sudden cliff-effects when historical events drop out of the window. EWMA provides continuous exponential smoothing for continuous metrics (like store traffic or logged overtime). For discrete integer events (like safety near-misses or customer grievances), EWMS provides recursive count decay without distortion.

Mathematical Formulation:

$$y_t = \alpha x_t + (1 - \alpha) y_{t-1}, \quad \alpha = \frac{2}{\text{span} + 1}$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.memory import calculate_ewma, calculate_ewms, scan_ewma_spans

# Continuous variable EWMA
traffic_ewma = calculate_ewma(
    series=labor_hours_series, # Input pandas Series or 1D array
    span=12,                   # Moving average decay span (e.g., 4, 12, 26, 52)
    half_life=None,            # Alternative half-life specification (e.g., 6.0, None)
    alpha=None,                # Direct smoothing factor override (0 < alpha <= 1, None)
    min_periods=1,             # Minimum required non-null observations
    adjust=True,               # Divide by decaying adjustment factor to mitigate early warmup bias
    ignore_na=False            # Whether to skip missing values in weight calculation
)

# Discrete event count EWMS (count data filtering)
incidents_ewms = calculate_ewms(
    series=incident_counts_series, # Discrete event counts (e.g., safety grievances, callouts)
    span=8,                        # Smoothing span
    min_periods=1,
    adjust=False                   # False preserves raw cumulative count scaling
)
EWMA continuous labor hours smoothing
Figure 6: EWMA continuous labor hours filtering across 104 weeks, mitigating short-term noise while tracking macro operational trends.
EWMS count data incident stream filtering
Figure 6b: EWMS count data filtering applied to discrete safety incidents, tracking cumulative risk momentum without boundary cliff-effects.

Feature 7: Double and Triple Exponential Smoothing (Holt-Winters)

Business Context: Enterprise workforce demand exhibits both long-term staffing growth (trend) and cyclical holiday or fiscal cycles (seasonality). Holt-Winters separates baseline level, slope trend, and seasonal components, enabling reliable multi-step headcount projections.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.memory import double_exponential_smoothing, triple_exponential_smoothing

# Triple exponential smoothing with additive seasonality and trend damping
smoothed_s, forecast_df, hw_res = triple_exponential_smoothing(
    series=workforce_demand_series, # 1D time-series Series or array
    seasonal_periods=52,            # Number of periods per seasonal cycle (e.g., 7 for daily, 52 for weekly)
    alpha=None,                     # Level smoothing factor (None for automated MLE fit)
    beta=None,                      # Trend smoothing factor (None for automated MLE fit)
    gamma=None,                     # Seasonal smoothing factor (None for automated MLE fit)
    trend='add',                    # Trend type: 'add', 'mul', or None
    seasonal='add',                 # Seasonality type: 'add', 'mul', or None
    damped_trend=True,              # Damp linear extrapolation to prevent runaway forecasts
    forecast_periods=12             # Number of out-of-sample periods to project forward
)
Holt-Winters double and triple exponential smoothing decomposition
Figure 7: Holt-Winters decomposition of a 3-year weekly workforce demand stream, isolating trend drift and 52-week seasonal hiring peaks.

Feature 8: Kaplan-Meier & Nelson-Aalen Hazard Embeddings

Business Context: Associate attrition is a classic time-to-event survival process. Associates who leave at 6 months exhibit fundamentally different risk dynamics than those with 5 years of tenure. This feature estimates non-parametric survival curves \(S(t)\) and cumulative hazards \(H(t)\) to embed historical tenure hazard directly into predictive feature matrices.

Mathematical Formulation:

$$\hat{S}(t) = \prod_{t_i \leq t} \left(1 - \frac{d_i}{n_i}\right), \quad \hat{H}(t) = \sum_{t_i \leq t} \frac{d_i}{n_i}$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.memory import fit_hazard_embeddings, HazardEmbeddingTransformer

# Requires: pip install 'people-analytics-toolkit[survival]'
# Fit non-parametric survival and hazard models on employee census
embedded_roster, hazard_models = fit_hazard_embeddings(
    roster_df=roster_df,             # Census DataFrame containing tenure and survival status
    tenure_col='tenure_weeks',       # Duration / tenure in position (e.g., 'tenure_weeks', 'tenure_months')
    event_col='turnover_event'       # Event indicator column (1 = voluntary termination, 0 = active / censored)
)
# Returns embedded_roster with: 'km_survival_prob', 'na_cum_hazard', 'hazard_quantile_bucket'
Kaplan-Meier survival and Nelson-Aalen hazard curves
Figure 8: Kaplan-Meier survival probability and Nelson-Aalen cumulative hazard curves across tenure cohorts, showing the 1-year and 3-year flight risk cliffs.

Feature 9: Configuration-Driven Rolling Metrics (Multi-Horizon Cohort Pipelines)

Business Context: Enterprise feature stores require consistent, reproducible feature pipelines that aggregate multiple operational signals across varying horizons (4, 12, 26, 52 weeks) without ad-hoc code duplication. This pipeline provides declarative multi-horizon rolling statistics.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.memory import compute_rolling_metrics, RollingMetricConfig, RollingMetricsPipeline

# Declarative multi-horizon cohort rolling feature generation
rolling_df = compute_rolling_metrics(
    df=store_telemetry_df,           # Panel DataFrame indexed by cohort and date
    date_col='fiscal_date',          # Chronological timestamp column
    metric_cols=['hours_worked', 'unplanned_absences'], # Numerical telemetry channels
    windows=[4, 12, 26, 52],         # Rolling horizons (e.g., 4-week, 12-week, 26-week, 52-week)
    aggregations=['mean', 'std', 'min', 'max', 'skew'], # Statistical aggregation functions
    group_col='department_id'        # Cohort grouping key (None for global un-grouped series)
)
Configuration-driven rolling metrics pipeline
Figure 9: Multi-horizon rolling metrics generated across department cohorts, showing rolling means, standard deviations, and kurtosis.

Feature 10: Singular Spectrum Analysis (SVD Smoothing & Lag-Free Baselines)

Business Context: Moving averages inevitably introduce phase lag, delaying the detection of operational shifts. Singular Spectrum Analysis (SSA) embeds the 1D time-series into a multidimensional trajectory Hankel matrix and applies SVD decomposition to reconstruct a clean, lag-free baseline trajectory.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.memory import singular_spectrum_analysis, SingularSpectrumAnalysis

# Extract lag-free signal baseline via Hankel SVD
ssa_results = singular_spectrum_analysis(
    series=daily_store_traffic,      # 1D series of operational activity or headcount
    window_length=14,                # Trajectory embedding window L (e.g., 7, 14, 30, 52; typically L <= N/2)
    top_k=3,                         # Leading singular value components to retain for reconstruction
    variance_threshold=None          # Optional cumulative variance cutoff (e.g., 0.90, 0.95, None)
)
# Returns dict: 'reconstructed', 'components', 'singular_values', 'explained_variance'
Singular Spectrum Analysis lag-free SVD reconstruction
Figure 10: Singular Spectrum Analysis (SSA) decomposing noisy daily store foot traffic into trend, harmonic seasonality, and high-frequency noise components.

Feature 11: Hierarchical Lag & Autoregressive Features

Business Context: Time-series forecasting for store-level turnover and headcount requires historical autoregressive anchors. Computing hierarchical lags within each organizational unit prevents cross-entity leakage across store boundaries.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.memory import compute_lag_features

# Generate group-aware autoregressive lag features
lagged_df = compute_lag_features(
    df=workforce_panel_df,           # Multi-entity panel DataFrame
    group_col='store_id',            # Entity grouping key (e.g., 'store_id', 'plant_id', 'department_id')
    time_col='fiscal_week',          # Monotonically increasing time integer or date
    value_cols=['turnover_count', 'overtime_hours'], # Metric columns to shift
    lags=[1, 2, 4, 12, 52]           # Autoregressive lag steps (1-wk, 2-wk, 1-mo, 1-qtr, 1-yr)
)
Hierarchical lag and autoregressive correlation
Figure 11: Hierarchical autoregressive lag structure across store clusters, showing seasonal 52-week correlation and short-term operational persistence.

Feature 12: Multi-Horizon Differencing & Growth Rates (YoY, MoM, WoW & % Change)

Business Context: Executive leadership monitors growth velocities: Year-over-Year (YoY), Month-over-Month (MoM), and Week-over-Week (WoW) percent changes. This function computes absolute deltas and relative growth rates across user-specified multi-horizons.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.memory import compute_differencing_features

# Multi-horizon growth velocities
diff_df = compute_differencing_features(
    df=workforce_panel_df,           # Panel DataFrame
    group_col='store_id',            # Cohort grouping key
    time_col='fiscal_week',          # Chronological index column
    value_cols=['headcount', 'labor_cost'], # Target metrics to evaluate
    periods=[1, 4, 12, 52],          # Horizons: 1 (WoW), 4 (MoM), 12 (QoQ), 52 (YoY)
    relative=True                    # True returns percentage growth rate: (x_t - x_{t-k}) / x_{t-k}
)
Multi-horizon growth rate velocities
Figure 12: Multi-horizon growth rate velocities (WoW, MoM, YoY) capturing acceleration and deceleration in departmental labor allocation.

Pillar 3: Capturing Chaos, Volatility & Regimes

Workplace systems experience volatility clustering: stable periods are punctuated by sudden waves of callouts, schedule churn, and resignations. When scheduling volatility spikes, absenteeism becomes contagious. Pillar 3 captures non-linear volatility, information entropy, and latent operational regimes.

Feature 13: Rolling Shannon Entropy (The "Chaos Index")

Business Context: When an employee's shift assignments swing erratically between morning, evening, night, and weekend rotations, their work-life predictability collapses. Rolling Shannon entropy quantifies scheduling chaos from categorical shift allocations. High entropy directly correlates with accelerated burnout and flight risk.

Mathematical Formulation:

$$H(X) = -\sum_{k=1}^{K} p_k \log_2(p_k), \quad H_{\text{norm}}(X) = \frac{H(X)}{\log_2(K)} \in [0, 1]$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.chaos import rolling_shannon_entropy, calculate_shannon_entropy

# Compute per-associate rolling scheduling entropy
entropy_series = rolling_shannon_entropy(
    df=shift_schedules_df,           # Shift assignment event DataFrame
    val_col='shift_type',            # Categorical role or shift assignment column (e.g., 'morning', 'night')
    window=14,                       # Rolling evaluation window in shifts/days (e.g., 7, 14, 28)
    group_col='employee_id',         # Associate grouping key
    base=2.0,                        # Logarithmic base (2.0 for shannons/bits, np.e for nats)
    normalize=True                   # True normalizes by log(K) so entropy ranges strictly in [0.0, 1.0]
)
Rolling Shannon Entropy Chaos Index
Figure 13: Rolling Shannon Entropy ("Chaos Index") over a 600-shift schedule stream. Sudden spikes reflect severe scheduling volatility prior to associate resignation.

Feature 14: GARCH(1,1) Dynamic Volatility Index

Business Context: Standard standard deviations assume constant variance (homoskedasticity). Real-world operational staffing variance clusters: high-variance days follow high-variance days. A GARCH(1,1) process models conditional variance as a dynamic function of past innovation shocks and past variance.

Mathematical Formulation:

$$\sigma_t^2 = \omega + \alpha \epsilon_{t-1}^2 + \beta \sigma_{t-1}^2, \quad \alpha + \beta < 1$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.chaos import fit_garch_volatility, rolling_garch_volatility, GARCHVolatilityModel

# Requires: pip install 'people-analytics-toolkit[timeseries]'
# Fit GARCH(1,1) model to unexpected scheduling variance residuals
cond_vol, garch_res = fit_garch_volatility(
    series=scheduling_variance_residuals, # 1D time-series of zero-mean operational shock residuals
    p=1,                                  # GARCH autoregressive lag order for past conditional variance
    q=1,                                  # ARCH moving average lag order for past squared innovation shocks
    mean='Constant'                       # Mean model formulation: 'Constant', 'Zero', or 'AR'
)
# Returns: cond_vol (conditional volatility series), garch_res (model summary with alpha, beta, omega)
GARCH(1,1) Dynamic Volatility Index
Figure 14: GARCH(1,1) conditional volatility tracking scheduling instability clusters across an operational distribution center.

Feature 15: Hawkes Process Intensity (The "Contagion" Feature)

Business Context: Unplanned absences are socially contagious. When two associates call out sick on a Monday, the remaining crew absorbs excessive strain, sharply increasing the probability that additional associates call out on Tuesday. A self-exciting Hawkes process models event arrivals where each historical event causes an immediate upward jump in intensity that decays exponentially over time.

Mathematical Formulation:

$$\lambda(t) = \mu + \sum_{t_i < t} \alpha \cdot \exp(-\beta(t - t_i))$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.chaos import HawkesProcessContagion

# Initialize Hawkes contagion model
hawkes = HawkesProcessContagion(
    baseline_mu=0.05,                # Baseline background arrival rate per day
    alpha=0.35,                      # Excitation jump magnitude per absence event
    beta=0.75                        # Exponential decay parameter governing speed of risk dissipation
)

# Fit maximum likelihood estimates from continuous event arrival timestamps
hawkes.fit(
    event_times=absence_arrival_days # 1D numpy array of logged event timestamps in days
)

# Evaluate dynamic absence contagion risk across continuous evaluation grid
contagion_intensity = hawkes.predict_intensity(
    query_times=evaluation_grid      # Array of query timestamps
)
Hawkes Process Absence Contagion Intensity
Figure 15: Hawkes process intensity curve tracking absence contagion spikes across a 120-day observation window.

Feature 16: Hidden Markov Model (HMM) Latent States

Business Context: Organizations transition through unobserved operational regimes: Stable Operations, Acute Crisis (e.g., peak holiday rush with high overtime and callouts), and Post-Crisis Recovery. A Gaussian Hidden Markov Model infers these hidden states from multivariate operational emissions.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.chaos import fit_workforce_hmm_regimes, WorkforceHMMRegimes

# Requires: pip install 'people-analytics-toolkit[timeseries]'
# Uncover latent operational regimes via multivariate Gaussian HMM
regime_df, trans_matrix, hmm_meta = fit_workforce_hmm_regimes(
    df=workforce_telemetry_df,       # Longitudinal multivariate operational telemetry
    emission_cols=['callout_rate', 'overtime_ratio', 'schedule_churn'], # Emission feature channels
    n_components=3,                  # Number of latent regimes (e.g., 2 for Stable/Crisis, 3 for Stable/Strained/Crisis)
    seed=42                          # Random seed for expectation-maximization initialization
)
# Returns: regime_df with 'latent_state' and posterior probabilities, trans_matrix (transition probabilities)
Hidden Markov Model Latent Operational Regimes
Figure 16: Gaussian HMM latent operational regimes inferred from callouts, overtime, and schedule churn, identifying critical organizational state transitions.

Feature 17: The Bradford Factor (Absenteeism Friction Score & Schedule Disruption)

Business Context: Operational scheduling is disrupted far more severely by frequent, short, unplanned absences than by a single prolonged medical leave. An associate absent for 1 day on 10 separate occasions causes widespread shift chaos, whereas an associate absent for 10 consecutive days causes a single manageable staffing adjustment. The Bradford Factor formalizes this asymmetry by squaring the number of spells.

Mathematical Formulation:

$$B = S^2 \times D$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.chaos import calculate_bradford_score, classify_bradford_risk

# Calculate Bradford friction score
b_score = calculate_bradford_score(
    instances=4,                     # Number of discrete, unplanned absence spells / episodes (S)
    total_days=10                    # Total cumulative days absent across all spells (D)
) # Formula: 4^2 * 10 = 160

# Classify operational risk tier
risk_tier = classify_bradford_risk(
    score=b_score                    # Returns: 'Low Risk' (<50), 'Moderate Concern' (50-200), 'Severe Disruption' (>200)
)
Bradford Factor Absenteeism Friction Score
Figure 17: Bradford Factor friction surface. Squaring absence spells penalizes intermittent disruptions far more heavily than sustained absences.

Pillar 4: Detecting Structural Anomalies, Density & Sequences

Workforce datasets contain subtle anomalies: compensation outliers masked by tenure, drift in departmental staffing structures, and abnormal career velocity profiles. Pillar 4 applies deep autoencoders, density estimation, trajectory alignment, and covariance-scaled distances.

Feature 18: Autoencoder Reconstruction Error (The "Weirdness" Index)

Business Context: Linear PCA misses non-linear interactions across multidimensional operational profiles (e.g., unusual ratios of managerial span, contract labor, and overtime). A deep bottleneck autoencoder compresses associate or store profiles into a low-dimensional manifold. High reconstruction loss signals structural anomalies.

Mathematical Formulation:

$$\mathcal{L}(x, \hat{x}) = \frac{1}{D} \sum_{d=1}^{D} (x_d - \hat{x}_d)^2$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.anomalies import AutoencoderWeirdnessDetector

# Requires: pip install 'people-analytics-toolkit[deeplearning]'
# Train deep bottleneck autoencoder on multidimensional store profiles
detector = AutoencoderWeirdnessDetector(
    hidden_dims=[32, 16, 8],         # Encoder layer sizes tapering to latent bottleneck (e.g., [64, 32, 16])
    learning_rate=0.005,             # Adam optimizer learning rate (e.g., 0.01, 0.005, 0.001)
    epochs=120,                      # Optimization epochs (e.g., 80, 120, 200)
    batch_size=32                    # Mini-batch gradient descent size (e.g., 16, 32, 64)
)

weirdness_scores = detector.fit_transform(
    X=store_feature_matrix           # Standardized continuous feature matrix of organizational profiles
)
Autoencoder Reconstruction Error Weirdness Index
Figure 18: Autoencoder reconstruction loss distribution isolating structural organizational anomalies on the right tail.

Feature 19: Local Outlier Factor (LOF via PyOD)

Business Context: Global outlier models fail when data clusters into regions of varying density (e.g., engineering salaries vs retail store wages). An executive earning $250k is normal; an entry-level clerk earning $140k in a low-cost region is a local anomaly. LOF evaluates an associate's local reachability density against their \(k\)-nearest neighbors.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.anomalies import fit_pyod_local_outlier_factor, PyODLocalOutlierFactor

# Requires: pip install 'people-analytics-toolkit[anomalies]'
# Fit density-based local outlier detection on compensation velocity
lof_scores, is_outlier, lof_meta = fit_pyod_local_outlier_factor(
    df=comp_velocity_df,             # Compensation and career progression DataFrame
    feature_cols=['salary_growth', 'tenure_years', 'performance_delta'], # Feature coordinates
    n_neighbors=20,                  # Neighborhood size k for reachability density (e.g., 15, 20, 30)
    contamination=0.04               # Expected prior proportion of true outliers (e.g., 0.01 to 0.05)
)
Local Outlier Factor LOF compensation velocity
Figure 19: PyOD Local Outlier Factor separating localized compensation outliers from legitimate tenure progression clusters.

Feature 20: Kullback-Leibler (KL) Divergence (The "Drift" Index)

Business Context: Retail stores and corporate divisions are budgeted for specific job role distributions (e.g., 60% associates, 25% specialists, 15% supervisors). When a business unit silently drifts into a top-heavy structure, payroll escalates. KL divergence quantifies the information-theoretic distance from the target benchmark.

Mathematical Formulation:

$$D_{\text{KL}}(P \parallel Q) = \sum_{k=1}^{K} P(k) \ln\left(\frac{P(k)}{Q(k) + \epsilon}\right)$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.anomalies import calculate_kl_divergence, calculate_jensen_shannon_divergence

# Calculate organizational staffing mix divergence from corporate benchmark
kl_drift = calculate_kl_divergence(
    p=observed_store_distribution,   # Empirical discrete probability distribution array (sums to 1.0)
    q=corporate_benchmark_profile,   # Target corporate distribution array (sums to 1.0)
    epsilon=1e-12                    # Numerical floor to prevent log(0) on empty job categories
)

# Symmetric Jensen-Shannon divergence bounded in [0, 1]
js_drift = calculate_jensen_shannon_divergence(
    p=observed_store_distribution,
    q=corporate_benchmark_profile
)
Kullback-Leibler Divergence Staffing Drift
Figure 20: KL Divergence ("Drift Index") across stores, highlighting units whose staffing mix has severely diverged from target headcount models.

Feature 21: Dynamic Time Warping (DTW) Distance

Business Context: Comparing seasonal staffing patterns across retail locations requires temporal flexibility. Store A might experience its holiday hiring rush in week 44, while Store B experiences it in week 46 due to regional weather. Euclidean distance misclassifies them as divergent; DTW non-linearly aligns the time axes.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.anomalies import compute_dtw_distance

# Non-linearly align seasonal curves
dtw_dist = compute_dtw_distance(
    s1=store_seasonal_curve,         # 1D series or array of local weekly staffing
    s2=corporate_benchmark_curve,    # 1D series of benchmark seasonal trajectory
    window=8,                        # Sakoe-Chiba warping window constraint (None for unconstrained warping)
    distance_metric='euclidean'      # Pointwise distance metric: 'euclidean', 'manhattan', 'sqeuclidean'
)
Dynamic Time Warping Seasonal Trajectory Alignment
Figure 21: Dynamic Time Warping warping path aligning phase-shifted seasonal staffing curves between retail stores.

Feature 22: Longest Common Subsequence (LCSS) & Edit Distance on Real Sequences (EDR)

Business Context: Career paths consist of ordered sequences of roles, certifications, and project assignments. An associate moving from [Sales Rep → Team Lead → Manager] follows a standard trajectory. LCSS and EDR measure similarity between real-valued career sequences under noise and gaps.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.anomalies import compute_lcss_distance, compute_edr_distance

# Compute sequence distance on career trajectory codes
lcss_distance = compute_lcss_distance(
    s1=associate_career_sequence,    # 1D array of career stage levels or weekly shift lengths
    s2=standard_career_template,     # Reference benchmark progression
    epsilon=1.0                      # Matching tolerance threshold for continuous grade equivalence
)

# Edit distance with sub-cost penalty
edr_distance = compute_edr_distance(
    s1=associate_career_sequence,
    s2=standard_career_template,
    epsilon=1.0
)
LCSS and EDR Career Sequence Trajectories
Figure 22: LCSS and EDR sequence similarity matrices comparing individual associate career progressions against standard talent pipelines.

Feature 50: Multivariate Mahalanobis Distance (Covariance-Scaled Anomaly Detection)

Business Context: Euclidean distance assumes uncorrelated, spherical features. In HCM, overtime, tenure, and compensation are highly correlated. An associate with high compensation and high tenure is normal; an associate with high compensation and low tenure is an extreme anomaly. Mahalanobis distance scales the distance by the inverse covariance matrix.

Mathematical Formulation:

$$D_M(x) = \sqrt{(x - \mu)^T \Sigma^{-1} (x - \mu)}$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.anomalies import calculate_mahalanobis_distance

# Compute covariance-scaled anomaly distance
mahalanobis_scores = calculate_mahalanobis_distance(
    df=workforce_kpi_df,             # DataFrame of multidimensional workforce metrics
    feature_cols=['overtime_hours', 'attrition_rate', 'span_of_control', 'compa_ratio'], # Feature vector
    robust=True                      # True fits Minimum Covariance Determinant (MCD) to prevent masking
)
Multivariate Mahalanobis Distance Chi-Squared Ellipses
Figure 50: Mahalanobis distance ellipses (\(\chi^2\) 95% and 99% contours) isolating multivariate outliers that escape univariate screening.

Pillar 5: Attribution, Interactions & Accounting

Executive HR leaders routinely confront aggregate macro movements: "Enterprise turnover rose from 14% to 18%." Did turnover rise because individual departments became worse (rate effect), or because the company expanded headcount in high-turnover retail divisions (mix effect)? Pillar 5 delivers exact rate-mix decomposition, non-linear SHAP interaction tensors, and robust correlation reports.

Feature 23: Log-Mean Decomposition (LMDI-I)

Business Context: Standard additive decompositions produce an unexplained residual interaction term. LMDI-I (Logarithmic Mean Divisia Index) utilizes logarithmic weights to achieve perfect zero-residual decomposition, isolating the exact percentage-point contribution of departmental turnover rates versus structural headcount shifts.

Mathematical Formulation:

$$L(a, b) = \frac{a - b}{\ln(a) - \ln(b)}, \quad \Delta R = \Delta R_{\text{rate}} + \Delta R_{\text{mix}} \quad (\text{Residual} \equiv 0)$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.attribution import lmdi_rate_mix_decomposition, logarithmic_mean

# Decompose aggregate turnover shift into rate vs mix effects
breakdown_df, total_effects = lmdi_rate_mix_decomposition(
    df=department_turnover_panel,    # DataFrame containing baseline and evaluation period statistics
    dept_col='department_name',      # Organizational division column
    h0_col='headcount_y0',           # Baseline headcount
    r0_col='turnover_rate_y0',       # Baseline departmental turnover rate
    h1_col='headcount_y1',           # Evaluation headcount
    r1_col='turnover_rate_y1'        # Evaluation departmental turnover rate
)
# Returns: breakdown_df (per-dept rate and mix deltas), total_effects ('delta_rate', 'delta_mix', 'delta_total')
Log-Mean Decomposition LMDI-I Rate Mix Decomposition
Figure 23: LMDI-I perfect decomposition proving whether organizational turnover spikes stem from worsening working conditions (rate effect) or structural hiring shifts (mix effect).

Feature 24: SHAP Interaction Values (2D Non-Linear Feature Interactions)

Business Context: Individual feature importances miss toxic synergies. For example, high overtime alone might increase turnover risk by +10%; a poor manager alone might increase risk by +15%. Combined, they might increase risk by +60%. SHAP interaction values decompose tree model predictions into main effects and pure pairwise non-linear synergies based on game-theoretic Shapley values.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.attribution import compute_shap_interaction_matrix

# Compute exact TreeSHAP pairwise interaction matrix
# Requires: pip install 'people-analytics-toolkit[explainability]' (on macOS: brew install libomp)
interaction_matrix = compute_shap_interaction_matrix(
    model=trained_tree_classifier,   # Trained scikit-learn, XGBoost, or LightGBM model
    X=features_df,                   # Feature matrix of employee covariates
    feature_names=None,              # Optional explicit list of feature names
    max_samples=500                  # Subsample limit for high-speed TreeSHAP evaluation
)
SHAP Pairwise Interaction Matrix
Figure 24: SHAP 2D interaction tensor exposing non-linear risk compounding between excessive overtime and low managerial quality.

Feature 25: Correlation Heatmap Table (Pearson, Spearman, Kendall with CIs & p-values)

Business Context: Presenting raw correlation numbers to leadership without statistical significance tests invites faulty policy decisions. This module computes comprehensive correlation reports including Pearson, Spearman rank, and Kendall tau statistics, complete with Fisher \(z\)-transformed 95% confidence intervals and exact \(p\)-values.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.attribution import correlation_heatmap_df

# Generate auditable correlation matrix with p-values and confidence bounds
corr_report_df = correlation_heatmap_df(
    df=workforce_analytics_df,       # Analysis DataFrame
    cols=['tenure', 'engagement', 'flight_risk', 'salary_growth'], # Columns to audit
    method='pearson',                # Correlation metric: 'pearson', 'spearman', 'kendall'
    alpha=0.05                       # Significance level for Fisher-z confidence intervals
)
Correlation Heatmap Table with Confidence Intervals
Figure 25: Correlation audit matrix with Fisher-z confidence intervals, distinguishing statistically validated behavioral drivers from random noise.

Pillar 6: Career Mobility, Succession & Behavioral Tipping Points

Career progression is not a continuous linear line; it operates through discrete graph transitions punctuated by stagnation cliffs. Associates stuck in grade beyond benchmark windows experience accelerated flight risk. Pillar 6 captures career mobility embeddings, zero-crossing SHAP inflection thresholds, and stagnation indices.

Feature 26: Movement Embeddings (Career Trajectory & Graph Embeddings)

Business Context: Standard talent analytics treats job titles as independent strings. However, "Senior Financial Analyst" shares latent competencies with "Operations Research Associate." By sampling career paths as transition walks through an organizational graph, Word2Vec/Skip-gram embeddings learn dense vector representations where cosine similarity directly mirrors internal mobility feasibility.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.mobility import CareerMovementEmbeddings, get_role_similarity_matrix

# Learn latent role vector representations from historical internal mobility logs
embedder = CareerMovementEmbeddings(
    vector_dim=16,                   # Latent embedding dimensionality (e.g., 8, 16, 32)
    window=3,                        # Context transition window
    min_count=1,                     # Minimum role transition observations required
    walk_length=10,                  # Length of random walks across organizational graph
    num_walks=20                     # Walks sampled per source job title
)

embedder.fit(
    transitions_df=career_moves_df,  # DataFrame with 'from_role' and 'to_role' columns
    source_col='from_role',
    target_col='to_role'
)

# Extract full cosine similarity matrix across enterprise job catalog
role_sim_df = get_role_similarity_matrix(embedder)
Career Movement Embeddings 2D Projection
Figure 26: 2D UMAP projection of learned Career Movement Embeddings, clustering interchangeable roles and uncovering non-traditional succession pathways.

Feature 27: SHAP Inflection Thresholds for Term Rates (Zero-Crossing Behavioral Tipping Points)

Business Context: Feature impact is rarely linear. Overtime might have zero impact on retention up to 5 hours/week, moderate impact from 5 to 12 hours, and exponential turnover risk beyond 12 hours. Spline-fitting SHAP dependence plots detects the exact mathematical zero-crossing inflection point (\(\text{SHAP}(x) = 0\)), providing actionable operational ceilings for managerial policies.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.attribution import extract_shap_inflection_thresholds

# Requires: pip install 'people-analytics-toolkit[explainability]'
# Detect exact zero-crossing tipping points from SHAP values
inflections, curves_df = extract_shap_inflection_thresholds(
    X=employee_covariates_df,        # Feature matrix
    shap_values=model_shap_values,   # Raw SHAP array from retention model
    n_grid_points=200,               # Interpolation grid resolution
    seed=42
)
# Returns: dict of inflection points per feature, and spline curves DataFrame
SHAP Inflection Threshold Zero Crossing Curves
Figure 27: Zero-crossing SHAP dependence curves identifying the exact overtime and tenure tipping points where retention risk flips positive.
Retained vs Terminated SHAP Contributions
Figure 27b: Aggregate SHAP value breakdown comparing retained associates against voluntary terminations across critical tipping features.

Feature 28: Time-in-Position (The Stagnation Index & Flight Risk Surge)

Business Context: An associate who has spent 36 months in an entry-level role whose median promotion time is 18 months experiences severe psychological stagnation. The Stagnation Index normalizes raw time-in-position against the empirical job family median, identifying associates at imminent risk of departure due to blocked mobility.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.mobility import calculate_time_in_position, calculate_stagnation_index

# Compute stagnation index relative to job family benchmarks
stagnation_df = calculate_stagnation_index(
    df=employee_roster_df,           # Employee history DataFrame
    current_date='2026-10-01',       # Evaluation anchor date string (YYYY-MM-DD)
    role_start_date_col='role_start', # Position inception date
    job_family_col='job_family',     # Cohort key for benchmark comparison
    benchmark_median_months=36.0     # Baseline benchmark months (or dict mapping job families)
)
Time-in-Position Stagnation Index Flight Risk Surge
Figure 28: Stagnation Index mapping associate duration against median job family promotion cycles, exposing the steep turnover surge beyond 1.5x benchmark.

Pillar 7: Compensation and Pay Equity Transformations

Raw base salary is uninformative in enterprise modeling because pay scales differ drastically across grades and regions. An engineer earning $120k may be severely underpaid, while a technician earning $80k may exceed their band maximum. Pillar 7 standardizes compensation into scale-free operational indices: Compa-Ratio and Range Penetration.

Feature 29: Compa-Ratio (Midpoint and Market Parity)

Business Context: Compa-Ratio (Comparative Ratio) measures where an employee's compensation sits relative to the established midpoint of their salary grade. A ratio of 1.00 indicates exact parity with market target; ratios below 0.85 indicate severe retention risk, while ratios above 1.15 indicate compensation compression.

Mathematical Formulation:

$$\text{Compa-Ratio} = \frac{\text{Base Salary}}{\text{Salary Band Midpoint}}$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.compensation import calculate_compa_ratio

# Calculate scale-free Compa-Ratio
compa_series = calculate_compa_ratio(
    salary=comp_df['base_salary'],   # Employee base salary Series or array
    midpoint=comp_df['band_midpoint'] # Salary band midpoint Series or constant float
)

Feature 30: Range Penetration (Band Progression & Red/Green Circled Status)

Business Context: While Compa-Ratio evaluates distance from the midpoint, Range Penetration evaluates progression through the entire pay band from minimum to maximum. Associates with Range Penetration > 1.0 are "Red-Circled" (overpaid relative to grade, ineligible for standard merit raises); associates < 0.0 are "Green-Circled" (underpaid, presenting legal and equity compliance risks).

Mathematical Formulation:

$$\text{Range Penetration} = \frac{\text{Base Salary} - \text{Band Minimum}}{\text{Band Maximum} - \text{Band Minimum}}$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.compensation import (
    calculate_range_penetration,
    classify_pay_band_status,
    is_red_circled,
    is_green_circled
)

# Compute Range Penetration
range_pen_series = calculate_range_penetration(
    salary=comp_df['base_salary'],   # Employee salary Series
    band_min=comp_df['band_min'],     # Band lower threshold Series
    band_max=comp_df['band_max']      # Band upper threshold Series
)

# Classify audit status: 'Green-Circled' (<0.0), 'Within-Band' [0.0, 1.0], 'Red-Circled' (>1.0)
audit_status = classify_pay_band_status(
    range_penetration=range_pen_series
)
Compensation Equity Bands, Compa-Ratio and Range Penetration Dashboard
Figure 29-30: 4-Panel Enterprise Compensation Parity Dashboard displaying salary band distributions, departmental Compa-Ratio vs parity, tenure compression, and audit status.

Pillar 8: Relational Dynamics, Structural Capital & ONA

Organizations operate as living collaboration graphs, not static tree hierarchies. An associate's formal job title rarely reflects their informal enterprise influence. When a critical cross-functional information broker departs, entire business units experience productivity shock. Pillar 8 computes rigorous Organizational Network Analysis (ONA) metrics, Ronald Burt's structural constraint, Euler magnitude, and span of control.

Feature 31: Betweenness Centrality (Cross-Functional Information Brokers)

Business Context: Information brokers sit on the shortest communication paths connecting disparate silos (e.g., bridging Engineering and Product). Betweenness centrality quantifies an associate's control over organizational information flow. Associates with high betweenness represent single points of organizational failure.

Mathematical Formulation:

$$C_B(v) = \sum_{s \neq v \neq t} \frac{\sigma_{st}(v)}{\sigma_{st}}$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.ona import calculate_betweenness_centrality

# Quantify cross-silo information brokerage
betweenness_dict = calculate_betweenness_centrality(
    G=enterprise_collaboration_graph, # NetworkX Graph or DiGraph of enterprise interactions
    normalized=True,                  # True normalizes by 2 / ((n-1)(n-2)) for cross-network comparability
    weight='weight'                   # Edge attribute holding tie strength / interaction frequency
)

Feature 32: Eigenvector Centrality (Informal Enterprise Hubs)

Business Context: Degree centrality merely counts an associate's direct connections. Eigenvector centrality weights connections by the centrality of the neighbors: being connected to five vice presidents yields vastly greater enterprise influence than being connected to twenty isolated interns.

Mathematical Formulation:

$$\lambda x_v = \sum_{u \in N(v)} A_{vu} x_u$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.ona import calculate_eigenvector_centrality

# Compute informal influence hub scores via power iteration
eigenvector_dict = calculate_eigenvector_centrality(
    G=enterprise_collaboration_graph,
    max_iter=1000,                    # Maximum power iteration convergence cycles (e.g., 500, 1000)
    tol=1e-6,                         # Convergence tolerance threshold
    weight='weight'                   # Edge interaction strength attribute
)

Feature 33: Ronald Burt's Network Constraint Index (Structural Holes & Redundancy)

Business Context: Ronald Burt's sociological theory of "structural holes" posits that associates whose contacts are redundant (all connected to one another in a dense clique) have high constraint, limiting their access to novel information. Associates spanning structural holes have low constraint, enabling creative cross-pollination and rapid innovation.

Mathematical Formulation:

$$C_i = \sum_{j} c_{ij}, \quad c_{ij} = \left(p_{ij} + \sum_{q \neq i \neq j} p_{iq} p_{qj}\right)^2$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.ona import calculate_burt_constraint

# Calculate Ronald Burt network constraint index
constraint_dict = calculate_burt_constraint(
    G=enterprise_collaboration_graph,
    nodes=None,                       # Subset of associate node IDs (None evaluates all nodes)
    weight='weight'                   # Dyadic interaction strength weight attribute
)

Feature 34: Euler Centrality & Leinster Metric Space Magnitude (Global Structural Capital)

Business Context: Local graph metrics miss global organizational diversity. Leinster metric space magnitude treats the enterprise graph as a metric space with similarity kernel \(Z_{uv} = \exp(-d(u,v))\). Solving the system \(Z \mathbf{w} = \mathbf{1}\) yields Euler weights \(\mathbf{w}\), quantifying each associate's non-redundant structural contribution to global organizational capital.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.ona import calculate_euler_centrality, compute_complete_ona_profile

# Compute Leinster magnitude Euler weights
euler_scores = calculate_euler_centrality(
    G=enterprise_collaboration_graph  # NetworkX enterprise interaction graph
)

# Or compute complete unified ONA feature profile in one vectorized call
full_ona_df = compute_complete_ona_profile(G=enterprise_collaboration_graph)
Enterprise Organizational Network Analysis 4-Panel Dashboard
Figure 31-34: Enterprise ONA Dashboard displaying network topology, betweenness broker scores, structural holes vs hub influence, Euler systemic capital, and departure contagion shock.

Feature 35: Span of Control (Managerial Load & Reorganization Flight Risk)

Business Context: Managerial span of control directly influences associate attrition. When restructuring leaves a manager with 22 direct reports instead of 7, 1-on-1 coaching evaporates, causing associate flight risk to surge. This module extracts hierarchical spans and evaluates reorganization shocks.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.ona import calculate_span_of_control, evaluate_span_reorganization_shock

# Extract reporting hierarchy spans
spans_df = calculate_span_of_control(
    df=hierarchy_roster_df,          # Census DataFrame
    employee_col='employee_id',      # Associate identifier column
    manager_col='manager_id'         # Supervisor identifier column
)

# Evaluate associate flight risk surge caused by restructuring span shocks
reorg_shock_df = evaluate_span_reorganization_shock(
    df_pre=pre_reorg_roster,         # Pre-restructuring reporting hierarchy
    df_post=post_reorg_roster,       # Post-restructuring reporting hierarchy
    employee_col='employee_id',
    manager_col='manager_id',
    base_flight_risk_col=None,       # Optional column containing baseline flight risk
    shock_sensitivity=0.15           # Multiplier per newly added direct report
)
Span of Control and Reorganization Shock Evaluation
Figure 35: Span of control distribution before and after organizational restructuring, tracking the managerial overload shock on associate flight risk.

Pillar 9: Algorithmic Equity, Pay Decomposition & Individual Fairness

People Analytics algorithms directly influence human livelihoods: compensation, hiring selection, performance ratings, and promotions. Deploying unconstrained models can automate historical bias and violate federal employment guidelines (EEOC, Title VII, EU AI Act). Pillar 9 implements rigorous econometric decomposition and algorithmic fairness constraints.

Feature 36: Oaxaca-Blinder Pay Gap Decomposition (Explained vs Unexplained Returns)

Business Context: Simple raw pay gap comparisons fail because they ignore differences in legitimate human capital endowments (tenure, role grade, education, performance). The econometric Oaxaca-Blinder technique decomposes the raw pay disparity into an Explained Portion (attributable to legitimate qualifications) and an Unexplained Portion (attributable to differential market returns / systemic bias).

Mathematical Formulation:

$$\bar{Y}_A - \bar{Y}_B = (\bar{X}_A - \bar{X}_B)\hat{\beta}^* + \left[ \bar{X}_A(\hat{\beta}_A - \hat{\beta}^*) + \bar{X}_B(\hat{\beta}^* - \hat{\beta}_B) \right]$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.equity import oaxaca_blinder_decomposition, OaxacaBlinderResult

# Perform econometric pay gap decomposition
ob_result = oaxaca_blinder_decomposition(
    df=equity_audit_df,              # Compensation audit census DataFrame
    outcome_col='base_salary',       # Remuneration outcome metric (e.g., 'base_salary', 'total_comp')
    group_col='protected_class',     # Binary sensitive group column (e.g., 'gender_code', 'ethnicity_code')
    feature_cols=[                   # Vector of legitimate human capital endowments
        'tenure_years',
        'job_level',
        'performance_rating',
        'education_years'
    ],
    group_majority=1,                # Label for majority / reference cohort (e.g., 1 or 'Male')
    group_minority=0,                # Label for comparison cohort (e.g., 0 or 'Female')
    reference_type='pooled'          # Reference model coefficients: 'majority', 'minority', 'pooled'
)
# Returns: ob_result.total_gap, ob_result.explained_gap, ob_result.unexplained_gap, ob_result.p_values

Feature 37: Demographic Parity Difference & Ratio (Macro Selection Equality)

Business Context: In talent acquisition and promotion algorithms, Demographic Parity requires that the positive selection rate be equal across protected demographic groups, enforcing the EEOC Four-Fifths (80%) Rule.

Mathematical Formulation:

$$\Delta_{\text{DP}} = |P(\hat{Y}=1 \mid A=1) - P(\hat{Y}=1 \mid A=0)| \leq \epsilon$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.equity import demographic_parity_difference, demographic_parity_ratio

# Evaluate selection parity differences across protected attributes
dp_diff = demographic_parity_difference(
    y_pred=promotion_decisions,      # Binary predicted promotion or interview decision (0 or 1)
    sensitive_features=gender_series # Sensitive demographic attribute Series
) # Ideal = 0.0

dp_ratio = demographic_parity_ratio(
    y_pred=promotion_decisions,
    sensitive_features=gender_series
) # Four-fifths rule check: ratio must be >= 0.80

Feature 38: Equalized Odds Constraint via Exponentiated Gradient Minimax Reduction

Business Context: Demographic Parity ignores true candidate performance, potentially penalizing qualified associates. Equalized Odds requires that True Positive Rates (TPR) and False Positive Rates (FPR) be equal across protected groups. The Exponentiated Gradient reduction solves a constrained minimax saddle-point game, optimizing predictive performance while bounding fairness violations.

Mathematical Formulation:

$$P(\hat{Y}=1 \mid A=1, Y=y) = P(\hat{Y}=1 \mid A=0, Y=y) \quad \forall y \in \{0, 1\}$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.equity import equalized_odds_difference, ExponentiatedGradientFairness
from sklearn.linear_model import LogisticRegression

# Audit equalized odds disparity
eq_diff = equalized_odds_difference(
    y_true=y_ground_truth,           # Ground truth outcome
    y_pred=y_predictions,            # Model predictions
    sensitive_features=gender_series # Protected category Series
)

# Train minimax fair model satisfying Equalized Odds constraints
fair_classifier = ExponentiatedGradientFairness(
    estimator=LogisticRegression(),  # Base scikit-learn classifier
    constraints='EqualizedOdds',     # Constraint: 'EqualizedOdds', 'DemographicParity', 'TruePositiveRateParity'
    eps=0.02,                        # Allowable constraint slack epsilon (e.g., 0.01, 0.02, 0.05)
    max_iter=50                      # Minimax saddle-point optimization iterations
)
fair_classifier.fit(X=X_train, y=y_train, sensitive_features=gender_train)

Feature 39: Individual Fairness & Local Predictive Consistency (Lipschitz Concordance)

Business Context: Group fairness metrics can be satisfied while still treating similar individuals unfairly (e.g., flipping decisions for two identical candidates). Dwork's Individual Fairness principle dictates: "Similar individuals must receive similar outcomes." This module calculates local Lipschitz neighborhood predictive concordance in job qualification space.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.equity import compute_individual_fairness_consistency

# Evaluate local individual predictive consistency across qualification neighbors
local_cons, global_cons, audit_meta = compute_individual_fairness_consistency(
    X=qualifications_feature_matrix, # Scaled continuous matrix of legitimate human capital attributes
    y_pred=model_predicted_probs,    # Model predicted probabilities or continuous scores
    n_neighbors=5,                   # Nearest neighbors k in qualification space (e.g., 3, 5, 10)
    distance_metric='euclidean'      # Metric in skill space: 'euclidean', 'manhattan', 'cosine'
)
# Returns: local consistency per associate, global consistency scalar, neighbor distance metadata
Algorithmic Equity and Fairness Audit Dashboard
Figure 36-39: 4-Panel Algorithmic Fairness Dashboard displaying Oaxaca-Blinder pay decomposition, exponentiated gradient parity convergence, group fairness enforcement, and individual consistency.

Pillar 10: Prescriptive Interventions: Causal Inference & Heterogeneous Uplift

Standard predictive models tell HR leaders who is likely to leave. They fail completely at answering what intervention will prevent them from leaving. Prescribing a retention bonus to an associate who would stay anyway wastes budget; prescribing it to a "sleeping dog" can trigger departure. Pillar 10 implements doubly robust causal inference, causal forests, uplift curves, and causal DAG discovery.

Feature 40: Augmented Inverse Propensity Weighting (AIPW Doubly Robust ATE)

Business Context: When evaluating voluntary workplace programs (such as mentorship or wellness stipends), associates self-select. High-performing associates enroll more frequently, creating massive confounding. AIPW combines a propensity model \(\hat{e}(X) = P(W=1 \mid X)\) with an outcome regression model \(\hat{\mu}(X, W)\). The estimator is "doubly robust": it remains completely unbiased if either the propensity model OR the outcome model is correctly specified.

Mathematical Formulation:

$$\hat{\tau}_{\text{AIPW}} = \frac{1}{N} \sum_{i=1}^{N} \left[ \hat{\mu}_1(X_i) - \hat{\mu}_0(X_i) + \frac{W_i(Y_i - \hat{\mu}_1(X_i))}{\hat{e}(X_i)} - \frac{(1 - W_i)(Y_i - \hat{\mu}_0(X_i))}{1 - \hat{e}(X_i)} \right]$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.prescriptive import aipw_estimator, AIPWResult

# Estimate true causal Average Treatment Effect (ATE) with double robustness
aipw_res = aipw_estimator(
    y=retention_outcomes,            # Binary or continuous outcome Series/array
    treatment=mentorship_enrolled,   # Binary policy intervention indicator (1 = enrolled, 0 = control)
    X=baseline_covariates_df,        # Baseline pre-treatment confounders (tenure, performance, role, compensation)
    propensity_model=None,           # Estimator for P(W=1|X) (None defaults to LogisticRegressionCV)
    outcome_model=None,              # Estimator for E[Y|X,W] (None defaults to RidgeCV)
    clip_range=(0.02, 0.98),         # Trimming bounds to prevent extreme propensity score weights
    n_splits=5,                      # K-fold cross-fitting splits to eliminate Donsker class bias
    random_state=42                  # Random seed for cross-fitting partition
)
# Returns: aipw_res.ate, aipw_res.se, aipw_res.p_value, aipw_res.ci_lower, aipw_res.ci_upper

Feature 41: Synthetic Difference-in-Differences (SDiD Panel Policy Evaluation)

Business Context: When an enterprise rolls out a new workplace policy (e.g., a 4-day workweek or scheduling app) in a single market, standard Difference-in-Differences (DiD) fails if the parallel trends assumption is violated. Synthetic DiD combines unit weights (creating a synthetic control twin) with time weights (re-weighting pre-intervention periods), delivering robust policy evaluation on panel data.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.prescriptive import synthetic_difference_in_differences

# Evaluate panel policy intervention effect
sdid_res = synthetic_difference_in_differences(
    df=workforce_panel_df,           # Longitudinal panel DataFrame across business units and time
    unit_col='business_unit',        # Unit / plant / market identifier column
    time_col='quarter_epoch',        # Time period index column
    outcome_col='turnover_rate',     # Continuous outcome metric (e.g., quarterly turnover rate)
    treated_units=['Office_North'],  # List of treated units receiving policy reform
    post_period_start=7,             # Starting epoch index of the post-intervention regime
    l2_regularization=1e-4           # L2 penalty on synthetic unit and time weight simplex estimation
)
# Returns dict: 'att' (Average Treatment Effect on the Treated), 'omega_weights', 'lambda_weights', 'p_value'

Feature 42: Causal Forest with Double Machine Learning (CausalForestDML & Honest Splitting)

Business Context: Treatment effects are heterogeneous: an intervention may decrease flight risk by -25% for high-tenure associates but increase flight risk by +10% for entry-level staff. Causal Forest with Double Machine Learning residualizes both treatment and outcome via cross-validated nuisance estimators, then builds an ensemble of honest causal trees to estimate individualized treatment effects (CATE).

Mathematical Formulation:

$$\tilde{Y} = Y - \hat{\mathbb{E}}[Y \mid X], \quad \tilde{W} = W - \hat{\mathbb{E}}[W \mid X], \quad \tau(x) = \arg\min_\tau \mathbb{E}\left[(\tilde{Y} - \tau(X)\tilde{W})^2\right]$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.prescriptive import CausalForestDML

# Train honest Causal Forest DML to estimate personalized treatment effects
cf_dml = CausalForestDML(
    n_estimators=50,                 # Number of honest causal trees in ensemble (e.g., 50, 100, 200)
    max_depth=4,                     # Maximum tree depth for heterogeneous treatment effect splits
    min_samples_leaf=10,             # Minimum sample size in treatment and control leaves
    subsample_ratio=0.7,             # Subsampling ratio for double-sample bagging
    propensity_model=None,           # Nuisance model for treatment residualization
    outcome_model=None,              # Nuisance model for outcome residualization
    random_state=42                  # Seed for reproducible honest sample partitioning
)

cf_dml.fit(
    Y=retention_1yr_series,          # 1-year retention outcome vector
    T=intervention_flag_series,      # Binary intervention treatment vector
    X=associate_covariates_df        # Predictor matrix for Conditional Average Treatment Effect (CATE)
)

individual_cate = cf_dml.predict_cate(X=associate_covariates_df)

Feature 43: Prescriptive Uplift Evaluation (Qini Curve, AUUC, and Doubly Robust MSE)

Business Context: ROC-AUC is useless for prescriptive targeting. An organization needs to prioritize the "Persuadables" (who only retain if treated) over "Sure Things" or "Lost Causes." The Qini curve sorts associates by predicted uplift and plots cumulative incremental gains over random targeting, quantifying the Area Under the Uplift Curve (AUUC).

API Usage & Full Parameter Specification:
from people_analytics_toolkit.prescriptive import qini_curve, area_under_uplift_curve, doubly_robust_uplift_eval

# Compute cumulative Qini targeting curve
qini_df = qini_curve(
    y_true=retention_outcomes,       # Ground truth outcome vector
    treatment=treatment_received,    # Binary treatment assignment
    uplift_scores=individual_cate,   # Predicted CATE uplift scores from Feature 42
    n_bins=10                        # Decile bins for cumulative uplift ranking (e.g., 10, 20)
)

auuc_metrics = area_under_uplift_curve(
    y_true=retention_outcomes,
    treatment=treatment_received,
    uplift_scores=individual_cate,
    n_bins=20
)

dr_eval = doubly_robust_uplift_eval(
    y_true=retention_outcomes,
    treatment=treatment_received,
    uplift_scores=individual_cate,
    X=baseline_covariates_df,
    clip_range=(0.01, 0.99)
)
Prescriptive Interventions Causal Inference and Uplift Dashboard
Figure 40-43: Prescriptive Uplift Dashboard displaying AIPW confounding adjustment, SDiD panel policy lift, CATE heterogeneity distributions, and Qini targeting efficiency.

Feature 44: Synthetic Control Weights (Synth-DiD Synthetic Twin & Causal Lift)

Business Context: When testing a retention initiative in a single flagship location, finding an identical control store is impossible. Synthetic Control solves a convex optimization problem to find a weighted combination of donor control stores that perfectly mirrors the pre-treatment trajectory of the pilot unit.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.prescriptive import build_synthetic_control_twin, compute_synthetic_control_weights

# Construct synthetic control twin for pilot retail store
synth_twin_df = build_synthetic_control_twin(
    df=retail_scheduling_panel,      # Panel DataFrame across stores and weeks
    unit_col='store_id',             # Unit identifier column
    time_col='fiscal_week',          # Longitudinal week identifier
    outcome_col='associate_turnover',# Target outcome metric
    pilot_unit='Store_01',           # Target pilot location
    post_period_start=16,            # Inception week of flexible scheduling deployment
    donor_pool=None,                 # Candidate control units (None uses all non-pilot stores)
    intercept=True,                  # Fit level-shift intercept to bridge baseline gaps
    l2_regularization=1e-4           # L2 ridge regularization penalty on donor weights
)
Synthetic Control Weights and Synthetic Twin Trajectory
Figure 44: Synthetic control twin trajectory tracking Pilot Store 01 pre-intervention and exposing the true causal turnover reduction post-intervention.

Feature 45: Directed Causal Edge Weights (DAG Discovery & Confounder Untangling)

Business Context: Observational correlations frequently confuse cause and effect. Does overtime cause turnover, or does impending turnover force remaining associates to log overtime? The NOTEARS (Non-combinatorial Optimization for Directed Acyclic Graphs) algorithm formulates DAG structure learning as a continuous optimization problem, eliminating circular feedback loops and discovering directed causal edges.

Mathematical Formulation:

$$\min_{W} \frac{1}{2n} \|X - XW\|_F^2 + \lambda_1 \|W\|_1 \quad \text{s.t.} \quad h(W) = \text{tr}\left(e^{W \circ W}\right) - d = 0$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.prescriptive import (
    compute_directed_causal_edge_weights,
    diagnose_causal_confounding,
    notears_linear
)

# Learn continuous causal DAG structure from observational data
dag_results = compute_directed_causal_edge_weights(
    df=workforce_observational_df,   # Observational DataFrame
    variables=[                      # Candidate causal graph nodes
        'overtime_hours',
        'manager_quality',
        'commute_distance',
        'flight_risk',
        'turnover_event'
    ],
    lambda1=0.03,                    # L1 sparsity penalty for graph edge pruning (e.g., 0.01, 0.03, 0.05)
    threshold=0.15,                  # Edge weight truncation threshold (e.g., 0.10, 0.15, 0.25)
    standardize=True                 # Standardize variables to zero mean and unit variance
)

# Diagnose whether correlation between two nodes is direct or confounded
confounding_diag = diagnose_causal_confounding(
    weights_df=dag_results['directed_weights_df'],
    cause_var='overtime_hours',
    effect_var='turnover_event',
    corr_matrix=dag_results['correlation_matrix']
)
Directed Causal Edge Weights NOTEARS DAG Discovery
Figure 45: Directed acyclic graph (DAG) learned via continuous NOTEARS optimization, separating direct causal drivers from spurious observational correlations.

Pillar 11: Continuous-Time Trajectories & Heterogeneous Graph Representation Learning

Real-world career progressions occur in continuous time: employees do not transition solely on fiscal year boundaries. Furthermore, associates are embedded in heterogeneous knowledge networks connecting people, skills, and cross-functional projects. Pillar 11 leverages continuous-time Markov chains, meta-path random walks, Metapath2Vec graph embeddings, and multi-head trajectory self-attention.

Feature 46: Continuous-Time Markov Chain (CTMC) Career Trajectories & Matrix Exponential

Business Context: Discrete Markov chains assume fixed time steps (e.g., annual transitions). Real career moves happen asynchronously. A Continuous-Time Markov Chain models transitions via an infinitesimal transition rate matrix \(\mathbf{Q}\). The continuous transition probability matrix for any arbitrary time horizon \(t\) is solved via the matrix exponential.

Mathematical Formulation:

$$\mathbf{P}(t) = \exp(\mathbf{Q} t) = \sum_{k=0}^{\infty} \frac{(\mathbf{Q} t)^k}{k!}, \quad \mathbb{E}[\text{Sojourn}_i] = -\frac{1}{q_{ii}}$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.trajectories import ContinuousTimeMarkovChain

# Initialize CTMC with career grade states
ctmc = ContinuousTimeMarkovChain(
    state_labels=['Associate', 'Senior', 'Staff', 'Director', 'Terminated']
)

# Fit transition intensity rate matrix Q from continuous event log
ctmc.fit_from_event_log(
    event_log=career_moves_df,       # Event log DataFrame
    employee_col='employee_id',      # Associate identifier
    state_col='role',                # Career stage / job rank column
    duration_col='duration_years',   # Continuous tenure in state
    censored_col='is_active'         # Censoring indicator (1 if ongoing, 0 if transitioned)
)

# Compute transition probability matrix P(t) for arbitrary forward horizons (e.g., 2.5 years)
P_matrix_2yr = ctmc.transition_probability_matrix(t=2.5)

# Calculate expected sojourn / holding times in each career grade before transition
sojourn_times = ctmc.expected_sojourn_times()

Feature 47: Meta-Path Biased Random Walks on Heterogeneous HCM Multigraphs

Business Context: An enterprise graph contains multiple entity types: Employees, Skills, and Projects. Unconstrained random walks produce meaningless traversals. Meta-path biased random walks force paths to follow specific organizational semantic templates (such as Employee → Project → Employee for collaboration affinity, or Employee → Skill → Employee for skill complementarity).

API Usage & Full Parameter Specification:
from people_analytics_toolkit.trajectories import generate_metapath_walks

# Generate semantic meta-path random walks over heterogeneous multigraph
sampled_walks = generate_metapath_walks(
    G=heterogeneous_hcm_multigraph,  # NetworkX Graph connecting employees, skills, and projects
    meta_paths=[                     # List of semantic node traversal templates
        ['employee', 'project', 'employee'],
        ['employee', 'skill', 'employee'],
        ['employee', 'project', 'skill', 'project', 'employee']
    ],
    walk_length=20,                  # Steps along each random walk (e.g., 10, 20, 40)
    num_walks=10,                    # Walk iterations initiated per node (e.g., 5, 10, 20)
    random_state=42                  # Seed for reproducible stochastic sampling
)

Feature 48: Metapath2Vec Graph Representation Learning & Complementarity

Business Context: Traditional team staffing matches associates by keyword search. Metapath2Vec applies heterogeneous Skip-gram embedding with negative sampling to the meta-path walk corpus generated in Feature 47. The resulting latent embeddings capture deep organizational complementarity, enabling algorithmic team formation and cross-functional staffing.

API Usage & Full Parameter Specification:
from people_analytics_toolkit.trajectories import Metapath2Vec

# Train heterogeneous graph embedding model
mp2v = Metapath2Vec(
    embedding_dim=32,                # Latent vector dimension (e.g., 16, 32, 64)
    walk_length=20,                  # Length of sampled walks
    num_walks=10,                    # Number of walks per node
    window_size=5,                   # Skip-gram context window size
    negative_rate=5.0,               # Negative sampling ratio
    random_state=42                  # Random seed
)

# Learn dense latent embeddings from meta-path walk corpus
node_embeddings = mp2v.fit_transform(
    walks=sampled_walks              # Tokenized trajectory corpus from Feature 47
)

Feature 49: Dynamic Entity & Trajectory Self-Attention Aggregator

Business Context: An associate's career history is an ordered sequence of roles, project milestones, and performance ratings. Traditional mean pooling treats all past milestones equally. A multi-head self-attention aggregator dynamically attends to the most formative career transitions, generating an adaptive employee representation for talent mobility models.

Mathematical Formulation:

$$\text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V$$
API Usage & Full Parameter Specification:
from people_analytics_toolkit.trajectories import DynamicEntitySelfAttention

# Initialize multi-head self-attention trajectory aggregator
attention_aggregator = DynamicEntitySelfAttention(
    n_heads=2,                       # Number of attention heads (e.g., 2, 4, 8)
    d_model=32,                      # Embedding dimensionality (must be divisible by n_heads)
    random_state=42                  # Weight initialization seed
)

# Aggregate sequential career stage embeddings into a unified associate vector
agg_vector, attn_weights = attention_aggregator.aggregate(
    embeddings_sequence=career_state_embeddings # 3D or 2D array of historical sequence vectors
)
# Returns: aggregated candidate representation vector and attention weight matrix
Continuous-Time Trajectories and Heterogeneous Graph Learning Dashboard
Figure 46-49: Continuous-Time Trajectories & Heterogeneous Graph Representation Dashboard displaying CTMC transition probabilities P(t), sojourn bottleneck durations, Metapath2Vec 2D projection, and self-attention candidate profile aggregation.
HR Analytics in Practice - Toby Flenderson
"There is no such thing as an appropriate joke. That's why it's called a joke." - Moving Human Resources away from legacy friction and exit questionnaires toward automated, mathematically robust People Analytics.

Conclusion & Production Deployment Recommendations

Transforming raw Human Capital Management telemetry into reliable, equitable, and causal signals requires dedicated mathematical architectures. By replacing ad-hoc scripts with people-analytics-toolkit, enterprise analytics teams achieve:

  • Leakage-Free Hierarchical Modeling: Empirical Bayes target shrinkage and grouped LOO z-scores eliminate target leakage in nested store and department hierarchies.
  • Continuous Memory Preservation: Fractional differencing and tier-weighted exponential decay model real-world employee event decay without destroying long-term organizational memory.
  • Operational Chaos & Volatility Tracking: Rolling Shannon entropy, GARCH(1,1), and Hawkes self-exciting processes provide early warning indicators for scheduling turbulence and absence contagion.
  • Econometric & Algorithmic Equity: Oaxaca-Blinder decomposition, equalized odds minimax reduction, and individual Lipschitz concordance provide end-to-end fairness compliance.
  • Prescriptive Causal Decision Support: AIPW doubly robust estimators, Synthetic DiD, and Causal Forest DML decouple correlational noise from actionable intervention policy.

For questions, issue tracking, or contributions to the package, visit the GitHub repository at github.com/sam-tritto/people-analytics-toolkit or inspect the release at pypi.org/project/people-analytics-toolkit/0.1.0.

More Projects