Φ(humanity): A Rigorous Ethical-Affective Objective Function

Formalizing Human Welfare for AI Systems

Working Paper Version: 3.0 (2026-07-26) Status: Under Development Authors: Research collaboration with Claude Opus 4.6 Project: Detective LLM - Information Gap Analysis System


Executive Summary

This paper presents Φ(humanity), a formal mathematical function for quantifying human welfare across eight dimensions: care, compassion, joy, purpose, empathy, love, protection, and truth. The function combines insights from welfare economics (Nash Social Welfare Function), capability theory (Sen, Nussbaum), care ethics (hooks, Gilligan), and standpoint epistemology (Collins) to create an inequality-sensitive, non-substitutable measure of societal well-being.

Key Innovation: The function operationalizes bell hooks' definition of love as "active extension for development" (hooks 2000), separating it from mere protection—a distinction absent in prior welfare formalizations. This enables precise detection of paternalistic regimes that provide material care and safeguarding but deny developmental support.

Application: Designed for Detective LLM, an investigative AI that detects information gaps in documentary records. Scarcity-weighted priorities derived from the function (§6.1) rank gaps threatening scarce welfare constructs, with special attention to marginalized communities (Collins 1990).

Limitations: The eight-construct taxonomy is Western-situated and requires adaptation for Ubuntu, Confucian, Buddhist, and Indigenous ethical frameworks. The function is intended for diagnostic use (gap prioritization), not optimization (Goodhart's Law concerns).


Table of Contents

  1. Theoretical Foundations
  2. The Eight Constructs
  3. Mathematical Formulation
  4. Philosophical Grounding
  5. Implementation
  6. Application to Information Gap Detection
  7. Limitations and Future Work
  8. Bibliography

Reading Guide

For philosophers: Focus on §2 (constructs), §4 (philosophical grounding), and the hooks/Berlin integration in §3.2.

For economists: Focus on §3 (mathematical formulation), especially the Nash SWF structure and Atkinson inequality sensitivity.

For AI researchers: Focus on §6 (application to gap detection) and the scarcity-weighted priority mechanism.

For activists/practitioners: Focus on §2.6-2.7 (love/protection distinction) and the worked examples in §3.4.


1. Theoretical Foundations

1.1 The Problem: Algorithmic Human Welfare

Can human flourishing be quantified? Amartya Sen's capability approach (Sen 1999) deliberately left this question open, arguing that agreement on capabilities need not presuppose agreement on relative weights. Martha Nussbaum (2000) enumerated ten central human capabilities but resisted algorithmic aggregation. The field of AI alignment (Russell 2019, Bostrom 2014) urgently needs formal welfare specifications but has struggled to bridge philosophical rigor with computational tractability.

Our position: Algorithmic welfare functions are necessary for AI systems making decisions affecting human well-being, but they must:

  1. Make normative commitments transparent (not hide them in "neutral" algorithms)
  2. Privilege marginalized voices (Collins 1990, standpoint epistemology)
  3. Prevent dimensional collapse (care without empathy = dystopia, Sen 1999)
  4. Remain incomplete and revisable (Sen's intentional incompleteness)

1.2 Why Nash Social Welfare?

Arrow's Impossibility Theorem (1951) proved that no aggregation rule can satisfy all desirable properties simultaneously. The Nash Social Welfare Function (Nash 1950, formalized by Kaneko & Nakamura 1979) makes a principled trade-off:

Nash SWF: Φ = ∏ᵢ xᵢ^θᵢ where θᵢ are weights summing to 1

Properties (Kaneko & Nakamura 1979), for the unmodified Nash SWF:

  • Pareto efficiency: Cannot improve one person without harming another
  • Symmetry: Treats individuals equally (before applying weights)
  • Scale invariance: Robust to utility rescaling
  • Multiplicative structure: Any dimension → 0 drives Φ → 0 (non-substitutability)

Caveat (v2.2): These properties hold for the unmodified Nash SWF and do not all survive our extensions. In particular, the divergence penalty (§3.4) makes Φ intentionally non-monotone in divergence-increasing directions — raising care while love stays suppressed deepens paternalism and can lower Φ — so global Pareto monotonicity is deliberately traded away. What we guarantee instead, locked in property tests (tests/inference/test_welfare_properties.py), is: Φ ∈ [0,1]; raising the most deprived construct never lowers Φ; uniform improvement across all constructs never lowers Φ.

This multiplicative structure encodes a Rawlsian maximin intuition (Rawls 1971): a society with perfect care but zero empathy has near-zero welfare.

1.3 Our Extension: Nash + Capabilities + Care Ethics + Ubuntu

We extend the Nash SWF with:

  1. Capability-theoretic inputs (Sen 1999, Nussbaum 2000): Eight dimensions of functioning
  2. Equity weighting (Rawls 1971, Atkinson 1970): Inverse-deprivation weights that dynamically prioritize the most deprived construct, replacing symmetric equal weights
  3. Community solidarity multiplier (Ubuntu philosophy): λ_L as meta-construct — welfare emerges from relational context, not individual metrics in isolation
  4. Recovery-aware floors: Constructs below hard floors receive community-mediated recovery potential. Key insight: care doesn't begin the uptick without community intervention
  5. Care ethics integration (hooks 2000, Gilligan 1982): Love as generative, not defensive
  6. Ubuntu synergy coupling (Alkire & Foster 2011): Renamed from synergy to make Ubuntu grounding explicit (η raised from 0.05 to 0.10)
  7. Divergence penalties: Explicitly penalize care-without-love and similar mismatches

2. The Eight Constructs

All constructs measured on [0,1] scale, using population-weighted Atkinson complements (1 - Aε) where Aε is the Atkinson inequality index with aversion parameter ε. This ensures the function responds to distribution, not just level.

2.1 Care (c)

Definition: Resource allocation meeting basic needs (Tronto 1993, Held 2006)

Measurable proxies:

  • Poverty rate inversion (1 - poverty_rate)
  • Access to healthcare, education, housing
  • Food security indicators

Atkinson adjustment: A society with high average resources but high inequality (top 10% have everything) gets low c score.

2.2 Compassion (κ)

Definition: Responsive support to acute distress (Nussbaum 1996)

Measurable proxies:

  • Emergency response capacity
  • Crisis intervention services
  • Disaster relief effectiveness

Distinguished from care: Compassion is reactive (emergency), care is proactive (baseline provision).

2.3 Joy (j)

Definition: Positive affect above subsistence (Csikszentmihalyi 1990, flow theory)

Measurable proxies:

  • Subjective well-being surveys
  • Life satisfaction indices
  • Positive affect measures (PANAS)

Non-utilitarian: Joy matters, but not as a substitute for other dimensions (multiplicative structure prevents joy-without-dignity scenarios).

2.4 Purpose (p)

Definition: Alignment of actions with chosen goals (Frankfurt 1971, hierarchical desires)

Measurable proxies:

  • Autonomy indices
  • Educational/vocational alignment
  • Reported sense of meaning (Steger et al. 2006)

Synergy with joy: Flow states (Csikszentmihalyi 1990) emerge from j × p coupling.

2.5 Empathy (ε)

Definition: Accuracy of perspective-taking across groups (Batson 1991)

Measurable proxies:

  • Intergroup contact quality (Pettigrew & Tropp 2006)
  • Discrimination indices (inverted)
  • Cross-cultural understanding measures

Critical for truth: Empathy × truth synergy—perspective-taking requires epistemic integrity.

2.6 Love (λ_L) — NEW CONSTRUCT

Definition: Active extension of self for another's growth and development (hooks 2000, p. 4)

bell hooks writes:

"The word 'love' is most often defined as a noun, yet... we would all love better if we used it as a verb... Love is as love does. Love is an act of will—namely, both an intention and an action. Will also implies choice. We do not have to love. We choose to love."

And critically:

"Definitions are vital starting points for the imagination. What we cannot imagine cannot come into being. A good definition marks our starting point and lets us know where we want to end up... The will to extend one's self for the purpose of nurturing one's own or another's spiritual growth." (hooks 2000, p. 4-5, emphasis added)

Distinguished from protection (λ_P): Love is generative (creating capacity for flourishing), protection is defensive (preventing harm).

Philosophical grounding: Isaiah Berlin's two concepts of liberty:

  • Negative freedom (λ_P): Absence of interference
  • Positive freedom (λ_L): Capability to self-actualize

hooks' insight: Safeguarding ≠ developmental support. A paternalistic state can protect citizens while denying them agency—high λ_P, low λ_L.

Measurable proxies:

  • Community capacity-building investments
  • Educational/developmental program accessibility
  • Mutual aid network strength
  • Mentorship/apprenticeship availability

2.7 Protection (λ_P)

Definition: Risk-weighted safeguarding of life and dignity from harm

Measurable proxies:

  • Violence rates (inverted)
  • Legal protection effectiveness
  • Safety net robustness
  • Physical security indices

Distinguished from love: Reactive harm prevention, not proactive growth support.

Why this split matters: Medical Apartheid (Washington 2006) documents both absent protection (experimental surgeries without consent) AND absent love (no developmental support for enslaved people as full humans). Conflating these obscures the dual violence.

2.8 Truth / Epistemic Integrity (ξ)

Definition: Accuracy and transparency of institutional records (Fricker 2007, epistemic injustice)

Measurable proxies:

  • Document suppression rates (inverted)
  • FOIA compliance
  • Testimonial credibility equity (Fricker 2007)
  • Contradiction rates in official records

Standpoint sensitivity: Collins (1990) argues official records reflect powerful actors' standpoint. High ξ requires contested records (community testimony, counter-narratives) to be preserved and accessible.


3. Mathematical Formulation

3.1 Core Function (Equity-Weighted, Community-Mediated)

Φ(humanity) = f(λ_L) · [∏ᵢ (x̃ᵢ)^wᵢ] · Ψ_ubuntu · (1 - Ψ_penalty) / Φ_max

where:

  • i ∈ {c, κ, j, p, ε, λ_L, λ_P, ξ} — the eight constructs
  • f(λ_L) = λ_L^γ (γ = 0.5) — Community solidarity multiplier. Ubuntu substrate: when community is low, all welfare degrades multiplicatively. (See §3.1.1)
  • x̃ᵢ — Recovery-aware effective inputs. Constructs above their hard floor pass through unchanged; below floor, recovery depends on trajectory (dx/dt) and community capacity (λ_L^0.5). As of v3.0 the community term is gated by damage breadth and stagnation persistence. (See §3.1.2, §3.1.4)
  • wᵢ = (1/x̃ᵢ) / Σⱼ(1/x̃ⱼ) — Equity-adjusted weights (inverse deprivation). Replaces symmetric Nash θ = 1/8 with Rawlsian maximin: weights shift dynamically toward the most deprived construct. (See §3.1.3)
  • Ψ_ubuntu — Ubuntu synergy term (η = 0.10, renamed from Ψ_synergy). Welfare gains emerge from relationships between paired constructs, not isolation. (See §3.3)
  • Ψ_penalty — Divergence penalty (μ = 0.15) for structural distortions; one-sided for the four primary pairs as of v2.2. (See §3.4)
  • Φ_max = 1 + 4η + η_curiosity = 1.48 — Normalization constant (v2.2): the analytic maximum of Ψ_ubuntu. Dividing by it makes a perfect society (all constructs = 1.0) score exactly Φ = 1.0.

Philosophical synthesis: The product ∏(x̃ᵢ)^wᵢ encodes Western capability theory (Sen 1999, Nussbaum 2000) — individual constructs with equity priority. The multiplier f(λ_L) and synergy Ψ_ubuntu encode Ubuntu relational philosophy ("umuntu ngumuntu ngabantu" — a person is a person through other persons). Neither term alone produces Φ.

Normalization: Φ ∈ [0,1]. In versions ≤2.1 the [0,1] claim was false — Ψ_ubuntu is unbounded above 1, so Φ(all ones) reached 1.48; v2.1.1 documented the [0, ~1.48] range but left the function unnormalized. v2.2 divides by Φ_max, restoring the bound: a perfect society scores exactly 1.0, and synergy gains are expressed within [0,1] rather than above it. Division by a positive constant preserves every ordering, so no v2.1 comparison flips from this step (the one-sided penalty of §3.4 does change some values).

3.1.1 Community Solidarity Multiplier

f(λ_L) = λ_L^γ     where γ = 0.5

The square root ensures diminishing returns — moving λ_L from 0.04 to 0.25 (√: 0.2→0.5) matters more than moving from 0.64 to 0.81 (√: 0.8→0.9). This encodes Ubuntu's claim that community is the substrate, not a bonus: welfare doesn't add community on top, it emerges from community.

λ_L f(λ_L) Interpretation
1.0 1.00 Full community solidarity — welfare undiminished
0.5 0.71 Moderate — 29% welfare degradation
0.25 0.50 Low — 50% degradation
0.04 0.20 Near-collapse — 80% degradation

Verification: Φ(λ_L=0.1, others 0.5) ≈ 0.068 vs. Φ(λ_L=0.8, others 0.5) ≈ 0.396 — an 83% reduction. This and the formula's other contractual properties (bounds, monotone directions, recovery caps, fail-loud inputs) are locked in tests/inference/test_welfare_properties.py; the earlier single-point criterion ("<50%") was trivially satisfiable and verified nothing.

Intentional quadruple influence of λ_L: Love (λ_L) appears in four distinct mechanisms within Φ, giving it outsized influence compared to other constructs:

  1. Community multiplier: f(λ_L) = λ_L^0.5 — multiplicative pre-factor on entire Φ
  2. Equity weight & Nash product: λ_L participates as one of eight constructs in ∏(x̃ᵢ^wᵢ)
  3. Recovery floors: Community capacity λ_L^0.5 · 0.5 governs recovery for all below-floor constructs
  4. Synergy + penalty: λ_L appears in two synergy pairs (c·λ_L and λ_L·ξ) and two penalty pairs ((c−λ_L)² and (λ_L−ξ)²)

This is not an accident — it is the core Ubuntu claim: community is the substrate of welfare, not one dimension among equals. The mathematical dominance of λ_L formalizes "umuntu ngumuntu ngabantu" (a person is a person through other persons). A sensitivity analysis (§7.5) quantifies the magnitude of this influence.

3.1.2 Recovery-Aware Floors

When a construct xᵢ falls below its hard floor (Nussbaum non-negotiable capability threshold), the effective input x̃ᵢ depends on two factors:

if xᵢ ≥ floorᵢ:
    x̃ᵢ = xᵢ                    (pass through)
else:
    dx̂/dt     = clip(dxᵢ/dt, −0.3, +0.3)    (v2.2: clipped trajectory input)
    trajectory = σ(10·(dx̂/dt) − 3)           (own recovery trend)
    community  = λ_L^0.5                      (community capacity)
    gated      = community · G_breadth · G_persist(stepsᵢ)   (v3.0 — see §3.1.4)
    recovery   = min(0.5, max(trajectory, gated·0.5))
    x̃ᵢ = xᵢ + (floorᵢ − xᵢ) · recovery

Hard floors per construct:

Construct Floor Rationale
c (care) 0.20 Basic needs non-negotiable (Nussbaum 2000)
κ (compassion) 0.20 Crisis response minimum
λ_P (protection) 0.20 Safety non-negotiable (Berlin 1969)
λ_L (love) 0.15 Community minimum
ξ (truth) 0.30 Epistemic integrity highest floor (Fricker 2007)
j (joy) 0.10 Lower but present
p (purpose) 0.10 Lower but present
ε (empathy) 0.10 Lower but present

Key insight: Community can partially compensate for stagnant trajectory. Care doesn't begin the uptick without community intervention. This produces three recovery signatures:

  1. Healing trajectory + strong community (dx/dt > 0, λ_L high) → rapid recovery toward floor
  2. Stagnant + strong community (dx/dt ≈ 0, λ_L high) → partial recovery (community compensates)
  3. Stagnant + no community (dx/dt ≈ 0, λ_L low) → near-raw value persists ("true collapse")

The sigmoid bias of −3.0 ensures that dx/dt=0 maps to ~0.047 (not 0.5), preventing trajectory from dominating community capacity for all λ_L < 1.0.

Gaming guard (v2.2): dx/dt must be a windowed regression slope over at least 3 observations and is clipped to [−0.3, +0.3] before the sigmoid; recovery credit is capped at 0.5, so a below-floor construct's effective value can never exceed the midpoint between its raw value and its floor. Without the cap, a single optimistic dx/dt = 1.0 lifted a fully collapsed construct (raw 0.01) to exactly its floor — an adversary could suppress the level metric and inject a positive trend. Reporting must always show (raw, effective) pairs and flag collapse whenever raw < floor/2, regardless of recovery credit.

Known discontinuity: the effective-value slope jumps from (1 − recovery) to 1 at each floor boundary. Numeric differentiation must not straddle a floor; the priority proxy of §6.1 avoids the issue by construction.

Lagged λ_L for circularity breaking: When computing λ_L's own recovery (λ_L is below its floor), using the current λ_L value creates a circular dependency: λ_L's effective value depends on community capacity, which is λ_L itself. The implementation resolves this by accepting an optional lam_L_prev parameter — the λ_L value from the previous timestep. When provided, λ_L's recovery uses lam_L_prev^0.5 instead of lam_L^0.5 for its own community capacity, while all other constructs continue using the current λ_L. This breaks the self-reference without affecting other constructs' recovery.

3.1.3 Equity-Adjusted Weights

wᵢ = (1/x̃ᵢ) / Σⱼ(1/x̃ⱼ)

Weights are computed on effective (recovery-adjusted) values x̃ᵢ, not raw metrics. This means once recovery floors lift a below-floor construct's effective value, the weight shifts less toward that construct — it has been partially compensated.

Properties:

  • Equal inputs (all x̃ᵢ = 0.5) → equal weights (wᵢ = 1/8)
  • One deprived construct (x̃_c = 0.1, others = 0.5) → care weight dominates (~0.42). Note this is stated in terms of the effective x̃_c. Feeding raw c = 0.1 (below its 0.20 floor) through v3.0 recovery yields x̃_c ≈ 0.124 and therefore w_c ≈ 0.365 — the recovery lift partially compensates the construct, so the weight shifts less than the raw value alone would imply.
  • Weights always sum to 1.0

This replaces the symmetric Nash θ = 1/8 with Rawlsian maximin: the formula automatically prioritizes whichever construct is most deprived, without manual weight tuning.

Known trade-off — partial substitutability: Equity weights create a degree of inter-construct substitutability absent in the original equal-weight Nash SWF. When care (c) drops, its weight increases, which amplifies the exponent on c but decreases weights on other constructs. This means improvements in the deprived construct can partially offset stagnation elsewhere — a weaker form of substitutability than additive models but stronger than pure multiplicative non-substitutability. We accept this trade-off because: (1) the multiplicative structure still enforces zero-construct → zero-Φ, preventing full substitution; (2) Rawlsian prioritization of the worst-off is more important than perfect non-substitutability; (3) recovery floors bound how far any construct can actually fall, limiting the practical scope of weight redistribution.

3.1.4 Damage-Profile-Gated Community Coupling (v3.0)

Introduced in v3.0 (ADR-032). Motivated by the resilience baseline audit (ADR-031), which found that v2.2 granted at least half the maximum recovery credit to 137 of 220 genuinely chronic trajectories at their nadir (false-optimism rate 0.6227 against a 0.20 bar). Across 299 below-floor construct-instances at chronic nadirs, the community-capacity branch won recovery_aware_input's max() 298 times: damaged constructs were being credited because an undamaged λ_L vouched for them. No choice of constant could withdraw that vouching, because the coupling was not conditioned on the record's own damage profile. v3.0 conditions it.

Two multiplicative gates apply to the community-capacity branch only. Neither touches the own-trajectory term — a construct that is genuinely improving still earns credit for improving.

Breadth gate — how widely damaged the profile is, computed once per profile from raw metrics:

shortfallᵢ    = max(0, (floorᵢ − xᵢ) / floorᵢ)      per-construct, normalized
mean_shortfall = (1/8) · Σᵢ shortfallᵢ
G_breadth      = (1 − mean_shortfall) ^ BREADTH_EXPONENT

A profile in broad, deep collapse cannot have its intact λ_L vouch at full strength; localized damage loses little. Worked values at the shipped exponent (6.0):

Profile G_breadth
all constructs 0.5 (nothing below floor) 1.000
single dip (c = 0.05, others 0.5) 0.554
two dips (c = 0.05, ξ = 0.10, others 0.5) 0.311
broad collapse (all 0.05, λ_L = 0.8) 0.006

Persistence gate — how long a construct has been stagnating below its floor:

G_persist(steps) = 1                                    if steps ≤ GRACE_STEPS
                 = exp(−(steps − GRACE_STEPS) / TAU)    otherwise

Community solidarity catalyzes recovery; if the recovery it promised never materializes, the formula stops crediting it. Worked values at the shipped constants (GRACE_STEPS = 0, TAU = 0.5):

steps G_persist
0 1.000000
1 0.135335
2 0.018316
3 0.00247875
5 4.53999 × 10⁻⁵
20 4.24835 × 10⁻¹⁸

Counter semantics. A construct's stagnation counter counts consecutive steps spent below its floor without improvement, where improvement means dx/dt strictly greater than IMPROVEMENT_EPS (0.005). Improvement, or rising to or above the floor, resets the counter to zero. update_stagnation_steps() is the single sanctioned maintenance path — the audit replay lens, the forecaster, and the investigation loop must all call it, so chronicity semantics cannot drift apart between instruments. The counter is optional state: when stagnation_steps is omitted the persistence gate stays fully open, which is why a static snapshot (e.g. the Research tab) is scored on breadth alone.

Shipped constants (selected under a pre-registered rule; see ADR-032):

Constant Value
BREADTH_EXPONENT 6.0
GRACE_STEPS 0
TAU 0.5
IMPROVEMENT_EPS 0.005
CONSTRUCT_FLOORS unchanged from v2.2

Backward compatibility. Setting breadth_exponent = 0.0 and passing no stagnation state makes G_breadth = G_persist = 1, recovering v2.2 semantics exactly. Profiles with every construct above its floor are bit-identical between v2.2 and v3.0, because the gates only reach the below-floor branch — §3.5 Case 1 is unchanged as a result.

Audit outcome at these constants (seed 20260725): resilience fraction R_f = 0.7769 (pre-registered band [0.731, 0.931]), false-optimism rate 0.0182 (absolute bar 0.25), collapse fraction 0.0961. The safety-net obligation ADR-031 left outstanding is discharged as measured by the audit lens; production call paths currently receive breadth gating only, since threading stagnation counters through the live investigation loop remains a recorded successor obligation.

3.2 Historical Note: Exponent Assignment

Superseded in v2.1. The v1.0/v2.0 formula used per-construct exponents αᵢ to encode concavity (diminishing returns for basic needs). In v2.1, these are replaced by equity-adjusted weights wᵢ = (1/x̃ᵢ)/Σ(1/x̃ⱼ) (see §3.1.3), which achieve the same Rawlsian maximin effect dynamically rather than through fixed exponents.

The original assignment is preserved here for reference:

Construct Group Exponent (v1.0-v2.0) Justification Citation
c, κ (basic needs) α = 0.7 Concave: first units matter most Rawls 1971, Atkinson 1970
j, p (experiential) α = 1.0 Linear: no inherent satiation Csikszentmihalyi 1990
ε, λ_L, λ_P (relational) α = 0.8 Mildly concave hooks 2000, Berlin 1969
ξ (epistemic) α = 1.0 Linear: truth is non-substitutable Fricker 2007

Why equity weights are superior: Fixed exponents cannot respond to the current state of deprivation. When care drops from 0.5 to 0.1, α=0.7 increases the gradient by a fixed ratio regardless of other constructs. Equity weights shift the entire weight vector toward care, amplifying the response in proportion to the actual imbalance across all eight dimensions.

3.3 Ubuntu Synergy Term (Relational Coupling)

Ψ_ubuntu = 1 + η · [√(c · λ_L) + √(κ · λ_P) + √(j · p) + √(ε · ξ)]
             + η_curiosity · √(λ_L · ξ)

where η = 0.10 (raised from 0.05 to reflect Ubuntu's centrality in the revised formula).

Geometric mean (√) penalizes imbalance within pairs more than linear coupling:

  • √(0.9 · 0.1) ≈ 0.30
  • √(0.5 · 0.5) = 0.50

Synergy pairs and justifications:

  1. Care × Love (c · λ_L): Material provision (c) + developmental extension (λ_L) = true flourishing (hooks 2000). Care without love = paternalistic control.

  2. Compassion × Protection (κ · λ_P): Emergency response (κ) + safeguarding infrastructure (λ_P) = effective crisis intervention. Compassion without protection = vulnerable support.

  3. Joy × Purpose (j · p): Positive affect (j) + goal-alignment (p) = flow states (Csikszentmihalyi 1990). Joy without purpose = hedonic but meaningless.

  4. Empathy × Truth (ε · ξ): Perspective-taking (ε) + epistemic integrity (ξ) = accurate cross-group understanding (Fricker 2007). Empathy without truth = manipulated solidarity.

  5. Love × Truth (λ_L · ξ): Investigative drive (curiosity) + epistemic integrity = genuine inquiry (hooks 2000, Fricker 2007). Truth without love = surveillance. Love without truth = willful ignorance.

3.3.1 Curiosity Cross-Pair (Love × Truth)

Ψ_curiosity = η_curiosity · √(λ_L · ξ)

where η_curiosity = 0.08.

Philosophical grounding:

  • hooks (2000): Love is "the will to extend one's self for the purpose of nurturing one's own or another's spiritual growth." Curiosity is that extension directed at understanding.
  • Fricker (2007): Epistemic integrity requires not just accurate records but the drive to investigate when records are incomplete.
  • The detective system's constitution: "honest analysis is an act of care, not aggression."

Why a cross-pair, not a 9th construct: Curiosity is not a capability in isolation. It is what emerges when love meets truth — the investigative impulse that makes someone follow a hunch into uncomfortable territory. It cannot exist without both:

  • Without love (λ_L suppressed by capitalism): curiosity collapses. You don't follow hunches when survival consumes capacity.
  • Without truth (ξ suppressed by institutions): curiosity has no target. You can't investigate what you can't see is missing.

Divergence detection: The penalty term now includes (λ_L - ξ)²:

  • Truth without love (high ξ, low λ_L) = surveillance. Institutional transparency serving control, not care.
  • Love without truth (high λ_L, low ξ) = willful ignorance. Community solidarity refusing uncomfortable facts.

Application to hypothesis scoring: Curiosity relevance = √(priority(λ_L) · priority(ξ)), where priority is the scarcity-weighted proxy of §6.1 — deliberately not the true gradient. The exact ∂Φ/∂x is sign-indefinite in divergence regions (§3.4), and a square root of a negative product is undefined; the proxy is positive by construction. Hypotheses at the love/truth intersection — the hunches that nobody follows because they're economically irrational — surface higher in the ranking.

3.4 Penalty Term (Divergence Punishment)

Ψ_penalty = μ · [max(0, c − λ_L)² + max(0, κ − λ_P)² + max(0, j − p)²
                 + max(0, ε − ξ)² + (λ_L − ξ)²] / 5

where μ = 0.15.

Five penalty pairs (4 primary one-sided + 1 curiosity cross-pair two-sided):

  1. max(0, c − λ_L)² — care-without-love = paternalism
  2. max(0, κ − λ_P)² — compassion-without-protection = vulnerable support
  3. max(0, j − p)² — joy-without-purpose = hedonic treadmill
  4. max(0, ε − ξ)² — empathy-without-truth = manipulated solidarity
  5. (λ_L − ξ)² — truth-without-love = surveillance; love-without-truth = willful ignorance

Why one-sided (v2.2): the theory names exactly one failure direction per primary pair. The v2.1 symmetric square also punished the reverse direction — a love-rich but resource-poor community was penalized for its poverty a second time through (c − λ_L)². The curiosity cross-pair stays two-sided because both of its directions are named failure modes.

Exposure asymmetry (deliberate): λ_L and ξ each appear in two of the five pairs; the other six constructs appear in one. Truth and love are this framework's protagonists, so divergences involving them are penalized more heavily by construction.

Monotonicity consequence: any penalty that grows with a construct's value makes Φ locally non-monotone in that direction — raising care in an already-paternalistic state (c > λ_L) can lower Φ. This is intended (more provision without more developmental support deepens the distortion), but it means Φ is not Pareto-monotone (§1.2 caveat) and must never be used as a naive optimization target (§7.2).

Intentional overlap between synergy and penalty pairs: The (c, λ_L) and (λ_L, ξ) pairs appear in both Ψ_ubuntu (synergy bonus via √(c·λ_L) and √(λ_L·ξ)) and Ψ_penalty (divergence punishment via (c−λ_L)² and (λ_L−ξ)²). This apparent double-counting is deliberate.

When c and λ_L diverge (e.g., c=0.9, λ_L=0.1 — a paternalistic regime): the synergy term already penalizes via diminished √(0.9·0.1) = 0.30 instead of √(0.5·0.5) = 0.50, and the penalty term additionally punishes via max(0, 0.8)² = 0.64. This double response is the formula's paternalism and white supremacy detection mechanism: systems that provide material care while denying developmental support are penalized through two independent channels, reflecting the dual violence documented in Washington (2006) — both the absence of love and the structural distortion of providing care without it. A single channel would under-weight this historically pervasive pattern.

Effect: A society with c=0.9, λ_L=0.1 incurs:

penalty contribution from max(0, c − λ_L)² = (0.8)² = 0.64
total penalty = 0.15 · (0.64 + other terms) / 5

This is small per pair but load-bearing when combined with missing synergy (√(0.9 · 0.1) ≈ 0.30 vs. √(0.5 · 0.5) = 0.50).

Proof that penalty cannot make Φ negative:

The factor (1 - Ψ_penalty) must remain positive for Φ to be interpretable. Since all constructs xᵢ ∈ [0,1], each squared divergence (xᵢ - xⱼ)² ≤ 1.0. The worst case (maximum divergence on all 5 pairs):

Ψ_penalty_max = μ · (1.0 + 1.0 + 1.0 + 1.0 + 1.0) / 5 = μ · 1.0 = 0.15

Therefore (1 - Ψ_penalty) ≥ 1 - 0.15 = 0.85 > 0 for all valid inputs. ∎

The implementation includes a max(0.0, phi) guard as a defensive measure, but it is mathematically unreachable given μ = 0.15. Increasing μ above 1.0 would break this guarantee and is not recommended.

3.5 Worked Examples

Case 1: Balanced Moderate Society (all = 0.5)

f(λ_L) = 0.5^0.5 = 0.707
All x̃ᵢ = 0.5 (above all floors → pass through)
G_breadth = 1.0 (nothing below floor → gates inert; identical to v2.2)
All wᵢ = 1/8 (equal deprivation → equal weights)
∏(x̃ᵢ^wᵢ) = 0.5^(8·1/8) = 0.5
Ψ_ubuntu = 1 + 0.10·(4·0.5) + 0.08·0.5 = 1.24
Ψ_penalty = 0 (all pairs balanced)
Φ = 0.707 · 0.5 · 1.24 · 1.0 / 1.48 ≈ 0.296

Case 2: Paternalistic Regime (c=0.9, λ_L=0.1, λ_P=0.9, others=0.5)

High material care + protection, but low developmental support:

f(λ_L) = 0.1^0.5 = 0.316          # community collapse
Equity weights: λ_L dominates (most deprived)
Ψ_ubuntu: √(0.9·0.1)=0.30 vs √(0.5·0.5)=0.50 — diminished synergy
Ψ_penalty = 0.15 · [max(0, 0.9−0.1)² + (0.1−0.5)²] / 5 = 0.024
            # (κ−λ_P) = (0.5−0.9) < 0 → one-sided pair contributes 0
G_breadth = 0.775                   # λ_L=0.1 is below its 0.15 floor
x̃_λ_L    ≈ 0.106                    # gated community recovery
Φ ≈ 0.069                           # 23% of the balanced case (0.296)

Interpretation: Paternalistic regime scores well below balanced society. The formula detects the structural distortion through three mechanisms: community multiplier collapse (f=0.316), equity weights shifting toward love, and divergence penalty firing on max(0, c−λ_L)².

Case 3: Care-Without-Love Dystopia (c=1, λ_L=0, others=0.5)

f(λ_L) = 0.01^0.5 ≈ 0.1           # near-zero (clamped at 0.01)
G_breadth = 0.449                   # λ_L collapsed below its floor
x̃_λ_L ≈ 0.017                       # gated recovery lifts only slightly
Φ ≈ 0.002                           # multiplicative collapse

Interpretation: Multiplicative structure structurally prevents dimensional collapse. Technical provision without developmental support = near-zero welfare. The community multiplier alone drives Φ toward zero when λ_L collapses. (Exactly zero is unreachable: inputs are clamped at 0.01 and recovery floors lift effective values slightly — see §4.3.)

Case 4: Community-Mediated Recovery (care drops to 0.05, community λ_L=0.6, dx_c/dt=0.0)

With both gates open (G_breadth = G_persist = 1, i.e. v2.2 semantics), the bare call gives:

x̃_c = recovery_aware_input(0.05, 0.20, 0.0, 0.6)
     = 0.05 + (0.20 − 0.05) · min(0.5, max(σ(−3), 0.6^0.5 · 0.5))
     = 0.05 + 0.15 · min(0.5, max(0.047, 0.387))
     = 0.05 + 0.15 · 0.387 ≈ 0.108   # below the v2.2 cap — unaffected

In situ under v3.0, the gates are not open: care sitting below its floor is itself damage, so the breadth gate fires. For the full profile (c = 0.05, all other constructs 0.6):

G_breadth = 0.554                       # one construct below floor
gated     = 0.6^0.5 · 0.554 · 1.0 = 0.429
x̃_c      = 0.05 + 0.15 · min(0.5, max(0.047, 0.429 · 0.5)) ≈ 0.082
Φ        ≈ 0.142                        # v2.2 scored 0.183 — a 23% reduction

after 20 consecutive stagnant steps (G_persist = 4.25 × 10⁻¹⁸):
x̃_c      ≈ 0.057                        # community credit fully withdrawn
Φ        ≈ 0.095

Interpretation: Care at 0.05 (far below its 0.20 floor) with a stagnant trajectory (dx/dt = 0) still recovers through community capacity rather than its own trend — the sigmoid at dx/dt = 0 produces only σ(−3) ≈ 0.047, negligible against the community term. This is the original insight and it survives: care doesn't begin the uptick without community intervention.

What v3.0 adds is a limit on how long, and how broadly, that intervention can be credited on faith. At one step of stagnation the community still lifts care to 0.082. After twenty steps of no improvement, the persistence gate has withdrawn the community credit entirely and the effective value falls to 0.057 — the same place it would sit with no community at all (λ_L → 0 gives 0.05 + 0.15 · 0.047 ≈ 0.057). A community that has been about to catalyse recovery for twenty steps is no longer evidence of recovery, and the formula stops treating it as such. This is the mechanism that took the audit's false-optimism rate from 0.6227 to 0.0182.


4. Philosophical Grounding

4.1 Care Ethics Integration

Held (2006) distinguishes caring labor (meeting needs) from caring relations (mutual recognition). We operationalize this as:

  • Care (c): The labor of provision
  • Love (λ_L): The relational extension for growth

hooks (2000) argues love requires both intention (will to extend self) and action (actual extension). Our λ_L measures the action component through observable developmental support.

4.2 Standpoint Epistemology

Collins (1990, p. 234):

"Oppressed groups are frequently placed in the situation of being listened to only if we frame our ideas in the language that is familiar to and comfortable for a dominant group. This requirement often changes the meaning of our ideas."

Our application: Construct measurements must include marginalized standpoints:

  • ξ (truth) requires contested records, not just official ones
  • ε (empathy) requires cross-group understanding, not just dominant-group perspective

Atkinson inequality-sensitive inputs ensure Φ responds to distribution across groups, not just averages.

4.3 Non-Substitutability (Sen's Capabilities)

Sen (1999, p. 76) argues capabilities are constitutively plural—freedom of speech cannot substitute for freedom from hunger. Our multiplicative structure enforces this: collapse in any dimension drives Φ toward zero. (Implementation honesty: inputs are clamped at 0.01 and recovery floors lift effective values slightly, so Φ reaches ~0.002 rather than exactly 0 — see Case 3, §3.5.)

Contrast with utilitarian additive models where high joy could compensate for absent protection. The Nash SWF prevents this dimensional collapse.

4.4 The Love/Protection Distinction (hooks + Berlin)

Berlin (1969): Two concepts of liberty:

  • Negative: Freedom from interference (protection)
  • Positive: Freedom to self-actualize (love as developmental support)

hooks (2000, p. 13):

"When we understand love as the will to nurture our own and another's spiritual growth, it becomes clear that we cannot claim to love if we are hurtful and abusive."

Our λ_L (love) measures active nurturing (positive freedom), λ_P (protection) measures safeguarding from harm (negative freedom). Both necessary; neither sufficient.

Paternalism detection: High c + λ_P but low λ_L reveals systems that provide materially and protect physically while denying agency—a pattern invisible in frameworks conflating love with protection.


5. Implementation

5.1 Python Reference Implementation

import math
from typing import Dict, Optional

ALL_CONSTRUCTS = ["c", "kappa", "j", "p", "eps", "lam_L", "lam_P", "xi"]

CONSTRUCT_FLOORS = {
    "c": 0.20, "kappa": 0.20, "lam_P": 0.20,
    "lam_L": 0.15, "xi": 0.30,
    "j": 0.10, "p": 0.10, "eps": 0.10,
}

ETA = 0.10
ETA_CURIOSITY = 0.08
MU = 0.15
GAMMA = 0.5

# v2.2: analytic max of the synergy term — divides Phi so all-ones scores 1.0
PHI_MAX = 1.0 + 4.0 * ETA + ETA_CURIOSITY

# v2.2: recovery guards — trajectory clipped, credit capped at midpoint
DX_DT_CLIP = 0.3
MAX_RECOVERY_CREDIT = 0.5

# v3.0 (ADR-032): damage-profile gates on the community-capacity branch
IMPROVEMENT_EPS = 0.005   # dx/dt above this resets the stagnation counter
BREADTH_EXPONENT = 6.0    # breadth suppression steepness
GRACE_STEPS = 0           # full-credit stagnation window
TAU = 0.5                 # post-grace decay constant

PRIMARY_PAIRS = [("c", "lam_L"), ("kappa", "lam_P"), ("j", "p"), ("eps", "xi")]
CURIOSITY_CROSS_PAIR = ("lam_L", "xi")

def sigmoid(x: float) -> float:
    return 1.0 / (1.0 + math.exp(-x))

def community_multiplier(lam_L: float, gamma: float = GAMMA) -> float:
    """f(λ_L) = λ_L^γ — Ubuntu substrate multiplier."""
    return max(0.01, lam_L) ** gamma

def recovery_aware_input(x_i, floor_i, dx_dt_i, lam_L, *,
                         gate_breadth_value=1.0, gate_persistence_value=1.0):
    """Recovery-aware effective input (v2.2: clipped + capped; v3.0: gated).

    Both gates default to 1.0, so an ungated call reproduces v2.2 exactly.
    The gates apply to the community branch only — never to own trajectory.
    """
    if x_i >= floor_i:
        return x_i
    dx_dt = max(-DX_DT_CLIP, min(DX_DT_CLIP, dx_dt_i))
    trajectory = sigmoid(10.0 * dx_dt - 3.0)
    community_capacity = max(0.01, lam_L) ** 0.5
    gated_community = (community_capacity
                       * gate_breadth_value * gate_persistence_value)
    recovery = min(MAX_RECOVERY_CREDIT,
                   max(trajectory, gated_community * 0.5))
    return x_i + (floor_i - x_i) * recovery


def update_stagnation_steps(prev, metrics, derivatives, floors):
    """Advance per-construct stagnation counters by one step (v3.0).

    Counts consecutive steps below floor without improvement. Improvement
    (dx/dt > IMPROVEMENT_EPS) or rising to/above the floor resets to zero.
    The single sanctioned counter-maintenance path — audit lens, forecaster,
    and investigation loop must all call this so semantics cannot diverge.
    """
    prev_counters = prev if prev is not None else {}
    out = {}
    for c in ALL_CONSTRUCTS:
        below = metrics[c] < floors[c]
        improving = derivatives.get(c, 0.0) > IMPROVEMENT_EPS
        out[c] = prev_counters.get(c, 0) + 1 if below and not improving else 0
    return out


def gate_breadth(metrics, floors, *, breadth_exponent=BREADTH_EXPONENT):
    """Damage-breadth gate in [0, 1] (v3.0) — profile-level, raw metrics."""
    total = 0.0
    for c in ALL_CONSTRUCTS:
        floor_c = floors[c]
        total += max(0.0, (floor_c - metrics[c]) / floor_c) if floor_c > 0 else 0.0
    mean_shortfall = total / len(ALL_CONSTRUCTS)
    return (1.0 - mean_shortfall) ** breadth_exponent


def gate_persistence(steps, *, grace_steps=GRACE_STEPS, tau=TAU):
    """Stagnation-persistence gate in (0, 1] (v3.0) — per-construct."""
    if steps < 0:
        raise ValueError(f"stagnation steps must be >= 0, got {steps}")
    if steps <= grace_steps:
        return 1.0
    return math.exp(-(steps - grace_steps) / tau)

def equity_weights(effective: Dict[str, float]) -> Dict[str, float]:
    """Inverse-deprivation weights: wᵢ = (1/x̃ᵢ) / Σⱼ(1/x̃ⱼ)."""
    inv = {c: 1.0 / max(0.01, v) for c, v in effective.items()}
    total = sum(inv.values())
    return {c: v / total for c, v in inv.items()}

def ubuntu_synergy(metrics: Dict[str, float]) -> float:
    """Ψ_ubuntu synergy on RAW metrics."""
    pairs = sum(math.sqrt(metrics[a] * metrics[b]) for a, b in PRIMARY_PAIRS)
    a, b = CURIOSITY_CROSS_PAIR
    return 1.0 + ETA * pairs + ETA_CURIOSITY * math.sqrt(metrics[a] * metrics[b])

def divergence_penalty(metrics: Dict[str, float]) -> float:
    """Ψ_penalty on RAW metrics (v2.2: primary pairs one-sided)."""
    sq = sum(max(0.0, metrics[a] - metrics[b]) ** 2 for a, b in PRIMARY_PAIRS)
    a, b = CURIOSITY_CROSS_PAIR
    sq += (metrics[a] - metrics[b]) ** 2
    return MU * sq / 5

def compute_phi(
    metrics: Dict[str, float],
    derivatives: Optional[Dict[str, float]] = None,
    lam_L_prev: Optional[float] = None,
    *,
    floors: Optional[Dict[str, float]] = None,
    stagnation_steps: Optional[Dict[str, int]] = None,
    breadth_exponent: float = BREADTH_EXPONENT,
    grace_steps: int = GRACE_STEPS,
    tau: float = TAU,
) -> float:
    """
    Compute Phi(humanity) — the full welfare function (v3.0).

    Phi = f(lam_L) * product(x_tilde_i ^ w_i) * Psi_ubuntu
          * (1 - Psi_penalty) / PHI_MAX

    v3.0: below-floor community recovery credit is gated by damage breadth
    (profile-level) and stagnation persistence (per-construct). v2.2
    semantics are recoverable exactly with breadth_exponent=0.0 and no
    stagnation state.

    v2.2: normalized to [0, 1]; all 8 constructs required — a missing
    measurement is an information gap, not "moderate welfare".

    Args:
        metrics: Dict mapping every construct symbol to a value in [0, 1].
        derivatives: Optional dict of dx/dt per construct. Defaults to 0.0.
        lam_L_prev: Previous timestep's λ_L, used to break the circular
            dependency in λ_L's own recovery (v2.1.1, audit fix #3).
            When None, falls back to current λ_L (backward-compatible).

    Raises:
        ValueError: If any construct is missing from `metrics`.
    """
    missing = [c for c in ALL_CONSTRUCTS if c not in metrics]
    if missing:
        raise ValueError(
            f"compute_phi requires all 8 constructs; missing: {missing}"
        )

    if derivatives is None:
        derivatives = {}

    lam_L_raw = max(0.01, metrics["lam_L"])
    f_lam = community_multiplier(lam_L_raw)

    # For λ_L's own recovery, use lagged value to break circularity
    lam_L_for_own_recovery = lam_L_prev if lam_L_prev is not None else lam_L_raw

    if floors is None:
        floors = CONSTRUCT_FLOORS

    # v3.0: damage-breadth gate is profile-level — computed once, raw metrics
    breadth = gate_breadth(metrics, floors, breadth_exponent=breadth_exponent)

    # Recovery-aware effective values
    effective: Dict[str, float] = {}
    for c in ALL_CONSTRUCTS:
        x_raw = max(0.01, metrics[c])
        community = lam_L_for_own_recovery if c == "lam_L" else lam_L_raw
        persistence = (
            gate_persistence(stagnation_steps[c],
                             grace_steps=grace_steps, tau=tau)
            if stagnation_steps is not None
            else 1.0
        )
        effective[c] = recovery_aware_input(
            x_raw, floors[c], derivatives.get(c, 0.0), community,
            gate_breadth_value=breadth,
            gate_persistence_value=persistence,
        )

    # Equity weights on effective values
    weights = equity_weights(effective)

    # Weighted geometric mean of effective values
    product = 1.0
    for c in ALL_CONSTRUCTS:
        product *= max(0.01, effective[c]) ** weights[c]

    # Synergy and penalty on RAW metrics
    synergy = ubuntu_synergy(metrics)
    penalty = divergence_penalty(metrics)

    phi = f_lam * product * synergy * (1.0 - penalty) / PHI_MAX
    return min(1.0, max(0.0, phi))

5.2 Constraint Layer (Rights-Based Floors)

Before applying Φ for decisions, enforce:

# Hard floors (non-negotiable minimums)
FLOORS = {
    'c': 0.2,       # Basic needs non-negotiable (Nussbaum 2000)
    'kappa': 0.2,   # Compassion floor
    'lam_P': 0.2,   # Protection non-negotiable
    'lam_L': 0.15,  # Love floor (slightly lower but essential)
    'xi': 0.3,      # Epistemic integrity minimum (Fricker 2007)
    # Others: 0.1
}

5.3 Measurement Protocol

For each construct, compute:

  1. Raw metric (e.g., poverty rate, violence rate)
  2. Atkinson inequality index Aε with ε=1.5 (moderate inequality aversion)
  3. Atkinson complement: xᵢ = 1 - Aε

This ensures Φ responds to equitable distribution, not just averages (Atkinson 1970).

Measurement caveat (v2.2): the Atkinson index requires a distribution of an individual-level cardinal variable. That exists for care (income, access microdata) but not for institution-level constructs like ξ (FOIA compliance) or κ (relief effectiveness). For those, fall back to group-disaggregated means with a between-group inequality term, and say so in reporting — the standpoint-sensitivity claim of §4.2 is only as real as the disaggregation behind it.


6. Application to Information Gap Detection

6.1 Gap Urgency via Scarcity-Weighted Priorities

Detective LLM detects information gaps (temporal, geographic, entity-level). A scarcity-weighted priority score per construct drives investigation ranking:

priority(xᵢ) = λ_L^0.5 · wᵢ / xᵢ

where wᵢ are the equity weights of §3.1.3. This is deliberately not the true ∂Φ/∂xᵢ. v2.1.1 (audit fix #4) computed the true gradient via central finite differences and clamped it to ≥ 0 on the claim that "welfare always improves with more of any construct" — but that claim is false: the divergence penalty (§3.4) makes ∂Φ/∂x negative in divergence regions, so the clamped gradient returns zero priority exactly where structural distortion lives — deprioritizing investigation where it matters most. Finite differences also straddle the derivative kinks at recovery floors (§3.1.2). v2.2 therefore reverts to an explicit proxy: always positive, rising with scarcity, honestly named. (The example values printed in versions ≤2.1.1, 1.43/0.16, were stale v1.0 exponent-formula outputs.)

def gap_urgency(
    gap: Gap,
    current_metrics: dict[str, float]
) -> float:
    """
    Urgency = Σ priority(xᵢ) for threatened constructs.

    Gaps threatening scarce constructs (high priority) = urgent.
    Gaps threatening abundant constructs (low priority) = less urgent.
    """
    threatened = gap.infer_threatened_constructs()

    priority_sum = sum(
        construct_priority(construct, current_metrics)
        for construct in threatened
    )

    return priority_sum * gap.confidence

Example (other constructs at 0.5):

  • Medical records gap (2010-2015) → threatens c (care), ξ (truth)
  • If c=0.1 (scarce): priority(c) ≈ 2.95 (HIGH urgency)
  • If c=0.9 (abundant): priority(c) ≈ 0.06 (LOW urgency)

6.2 Construct Inference Heuristics

Text → Construct mapping:

Text Pattern Constructs Example
"experimental surgery without consent" λ_P, ξ Washington 2006 (Medical Apartheid)
"mutual aid networks excluded" λ_L, c Hayes & Kaba 2023
"redacted oversight documents" ξ Fricker 2007 (testimonial injustice)
"community healing spaces defunded" λ_L, κ hooks 2000

6.3 Standpoint-Sensitive Prioritization

When official records are silent but community testimony exists:

  • ξ (truth) is LOW (suppression detected)
  • Gap urgency is HIGH (official/community divergence)

This operationalizes Collins' (1990) standpoint epistemology: contested absences are prioritized.


7. Limitations and Future Work

7.1 Acknowledged Limitations

  1. Proxy Validity: Function is only as good as construct measurements. Poor proxies → poor optimization (Goodhart's Law).

  2. Weight Setting is Political: v2.1+ replaced equal weights with inverse-deprivation equity weights (§3.1.3) — itself a normative commitment (Rawlsian maximin) that a community may reject in favor of equal weights or its own priorities. No formula bypasses normative debate (Sen 1999).

  3. Cultural Specificity: The 8-construct taxonomy is Western-situated. Ubuntu (collective humanity), Confucian (filial piety), Buddhist (non-attachment), and Indigenous (land reciprocity) frameworks require different construct spaces, not just weight tuning.

  4. Measurement Challenges: Measuring "love as developmental extension" is harder than measuring poverty. Requires:

    • Community capacity-building investments (observable)
    • Mutual aid network strength (network analysis)
    • Educational accessibility (administrative data)

7.2 Goodhart's Law Resistance

"When a measure becomes a target, it ceases to be a good measure." (Goodhart 1975)

Our mitigation:

  • Diagnostic use only: Φ prioritizes gaps, doesn't optimize policies directly
  • Multi-proxy triangulation: Each construct measured via 3-5 independent proxies
  • Red-team auditing: Adversarial testing for gaming vulnerabilities
  • Constitutional constraints: Hard floors + rate-of-change limits prevent sacrifice trades
  • Known non-monotone directions: §3.4 documents where improving a construct lowers Φ. An optimizer would exploit exactly these directions (e.g., suppressing ξ to shrink divergence penalties) — a further reason Φ must stay diagnostic

7.3 Open Questions

  1. Time-Varying Exponents: Should αᵢ adapt based on global scarcity? (If global c is high, shift to αc=1.0 for linear returns)

  2. Non-Human Life: How to integrate ecosystem integrity E? As a floor constraint (E ≥ 0.3 required for Φ calculation) or as a 9th construct?

  3. Optimization vs. Diagnostic: Should Φ ever be maximized, or only used to detect welfare drops? Maximization invites gaming (Goodhart); diagnostic use is safer.

  4. Cross-Cultural Validation: How do Ubuntu, Confucian, Buddhist, and Indigenous communities respond to this formulation? Requires co-design, not imposition.

7.4 Future Directions

  1. Empirical Calibration: Test on historical datasets (UNDP HDI, Alkire-Foster MPI) to validate scarcity-priority rankings

  2. Cultural Extensions: Partner with non-Western scholars to develop alternative construct spaces

  3. Constitutional AI Integration: Use scarcity-weighted priorities (§6.1) to shape LLM training (preference pairs weighted by welfare impact). Warning: this is optimization pressure against Φ and is in tension with the diagnostic-only stance of §7.2. Before any such run: freeze the metrics used for weighting, and red-team the non-monotone directions of §3.4 — they are exactly what a trained model would exploit

  4. Longitudinal Studies: Track Φ evolution over time to detect welfare trajectory changes

7.5 Parameter Sensitivity Analysis

The following table shows the effect of ±10% parameter changes on Φ at the balanced baseline (all constructs = 0.5). Sensitivity is measured as |ΔΦ/Φ_baseline| for a ±10% parameter shift. (Computed at v2.1.1, Φ_baseline ≈ 0.438; under v2.2 the baseline is Φ ≈ 0.296 after Φ_max normalization. Relative sensitivities are unchanged for γ and μ, which do not enter Φ_max; for η and η_curiosity the fixed normalizer partially offsets the synergy gain, so their table entries are upper bounds. v3.0 leaves this baseline untouched at Φ = 0.296220 — the gates reach only the below-floor branch, so the whole table carries over.)

Parameter Default +10% Value ΔΦ/Φ (%) Interpretation
γ (community exponent) 0.50 0.55 −3.4% Higher γ penalizes low community more; most sensitive parameter
η (synergy coupling) 0.10 0.11 +0.9% Modest: synergy is a multiplicative bonus ≥ 1.0
μ (penalty weight) 0.15 0.165 ≤ 0.0% Zero at balanced baseline (all pairs equal); up to −1.5% at maximum divergence
η_curiosity 0.08 0.088 +0.4% Smallest: single cross-pair term
Sigmoid bias −3.0 −3.3 < 0.1% Only affects below-floor constructs; negligible at baseline
Floor (care) 0.20 0.22 < 0.1% Only affects constructs below floor; irrelevant at baseline
BREADTH_EXPONENT (v3.0) 6.0 6.6 0.0% Gates the below-floor community branch only; exactly inert at baseline
TAU (v3.0) 0.5 0.55 0.0% Same: no construct is below floor, so no counter is running

Key findings:

  1. γ dominates: The community multiplier exponent is the most sensitive parameter. This reflects the Ubuntu design: λ_L's influence is intentionally outsized (see §3.1.1).
  2. η and μ are moderate: Synergy and penalty parameters have bounded effects because Ψ_ubuntu ∈ [1.0, ~1.48] and Ψ_penalty ∈ [0, 0.15].
  3. Floors and sigmoid bias matter only in crisis: These parameters are irrelevant at the balanced baseline but become dominant when constructs drop below their floors.
  4. λ_L construct value: Not a parameter but worth noting — a ±10% change in λ_L (from 0.5 to 0.45/0.55) produces ~8% change in Φ through the four channels documented in §3.1.1. This is roughly 2× the sensitivity of any other single construct.

7.6 Φ as Static Snapshot

Φ(humanity) is a static function, not a dynamic model. Given a set of construct measurements {xᵢ} and optional derivatives {dxᵢ/dt}, Φ returns a single scalar welfare score. It does not model how welfare evolves over time — that requires the PhiTrajectoryForecaster (src/forecasting/), which takes a time series of Φ values and predicts future trajectories.

The derivatives {dxᵢ/dt} in the recovery floor mechanism (§3.1.2) are externally provided rates of change, not computed by Φ itself. They inform recovery potential at a single point in time. Temporal dynamics — trend extrapolation, trajectory prediction, intervention modeling — are the domain of the forecaster, not the welfare function.

Implication for gap detection: Φ gradients (§6.1) prioritize gaps at a snapshot in time. To detect emerging welfare threats (constructs trending toward collapse), combine Φ gradient ranking with trajectory urgency from the forecaster (ADR-010).


8. Bibliography

Welfare Economics and Social Choice

Arrow, K. J. (1951). Social Choice and Individual Values. Yale University Press.

Atkinson, A. B. (1970). On the measurement of inequality. Journal of Economic Theory, 2(3), 244-263.

Goodhart, C. (1975). Problems of monetary management: The UK experience. Papers in Monetary Economics (Vol. 1). Reserve Bank of Australia.

Harsanyi, J. C. (1955). Cardinal welfare, individualistic ethics, and interpersonal comparisons of utility. Journal of Political Economy, 63(4), 309-321.

Kaneko, M., & Nakamura, K. (1979). The Nash social welfare function. Econometrica, 47(2), 423-435.

Nash, J. (1950). The bargaining problem. Econometrica, 18(2), 155-162.

Rawls, J. (1971). A Theory of Justice. Harvard University Press.

Capability Approach

Alkire, S., & Foster, J. (2011). Counting and multidimensional poverty measurement. Journal of Public Economics, 95(7-8), 476-487.

Nussbaum, M. (1996). Compassion: The basic social emotion. Social Philosophy and Policy, 13(1), 27-58.

Nussbaum, M. (2000). Women and Human Development: The Capabilities Approach. Cambridge University Press.

Sen, A. (1999). Development as Freedom. Oxford University Press.

Steger, M. F., Frazier, P., Oishi, S., & Kaler, M. (2006). The meaning in life questionnaire: Assessing the presence of and search for meaning in life. Journal of Counseling Psychology, 53(1), 80.

Care Ethics and Feminist Theory

Collins, P. H. (1990). Black Feminist Thought: Knowledge, Consciousness, and the Politics of Empowerment. Routledge.

Gilligan, C. (1982). In a Different Voice: Psychological Theory and Women's Development. Harvard University Press.

Held, V. (2006). The Ethics of Care: Personal, Political, and Global. Oxford University Press.

hooks, b. (2000). All About Love: New Visions. William Morrow.

Tronto, J. C. (1993). Moral Boundaries: A Political Argument for an Ethic of Care. Routledge.

Epistemology and Social Justice

Berlin, I. (1969). Two concepts of liberty. In Four Essays on Liberty (pp. 118-172). Oxford University Press.

Fricker, M. (2007). Epistemic Injustice: Power and the Ethics of Knowing. Oxford University Press.

Washington, H. A. (2006). Medical Apartheid: The Dark History of Medical Experimentation on Black Americans from Colonial Times to the Present. Doubleday.

AI Alignment and Constitutional AI

Anthropic. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv preprint arXiv:2212.08073.

Bostrom, N. (2014). Superintelligence: Paths, Dangers, Strategies. Oxford University Press.

Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. arXiv preprint arXiv:1706.03741.

Russell, S. (2019). Human Compatible: Artificial Intelligence and the Problem of Control. Viking.

Psychology and Well-Being

Batson, C. D. (1991). The Altruism Question: Toward a Social-Psychological Answer. Psychology Press.

Csikszentmihalyi, M. (1990). Flow: The Psychology of Optimal Experience. Harper & Row.

Frankfurt, H. G. (1971). Freedom of the will and the concept of a person. Journal of Philosophy, 68(1), 5-20.

Pettigrew, T. F., & Tropp, L. R. (2006). A meta-analytic test of intergroup contact theory. Journal of Personality and Social Psychology, 90(5), 751.

Application Context

Hayes, K., & Kaba, M. (2023). Let This Radicalize You: Organizing and the Revolution of Reciprocal Care. Haymarket Books.

Wilkerson, I. (2020). Caste: The Origins of Our Discontents. Random House.


Changelog

Version 3.0 (2026-07-26): Damage-Profile-Gated Community Coupling

Calibration release answering the resilience baseline audit (ADR-031), which measured the v2.2 formula's implied resilience distribution against Bonanno's 83.1% baseline. v2.2 cleared the resilience horn (R_f = 0.7802, inside the pre-registered band [0.731, 0.931]) but failed minority protection by roughly 3×: false-optimism rate 0.6227 against a 0.20 bar, meaning 137 of 220 genuinely chronic trajectories were granted at least half the maximum recovery credit at their nadir. A corrected stratified sweep found 0 of 210 grid points satisfying both criteria, establishing that the failure could not be discharged parametrically. The mechanism was structural: across 299 below-floor construct-instances at chronic nadirs, the community-capacity branch won recovery_aware_input's max() 298 times — damaged constructs were credited because an undamaged λ_L vouched for them, and no constant could withdraw a vouching that was never conditioned on the record's own damage profile.

The change (§3.1.4): two multiplicative gates on the community-capacity branch only, never on own trajectory.

  1. Breadth gate — (1 − mean_shortfall) ** BREADTH_EXPONENT, where mean_shortfall averages per-construct normalized shortfall over all 8 constructs. Always active. A profile in broad, deep collapse cannot have its intact λ_L vouch at full strength.
  2. Persistence gate — full credit inside a grace window, then exp(−(steps − grace) / TAU). Opt-in via per-construct stagnation counters maintained by update_stagnation_steps(). Community catalyzes recovery; if the recovery never materializes, the formula stops crediting it.

Shipped constants (chosen under a pre-registered adoption rule, re-registered after a measurement defect was found and corrected): BREADTH_EXPONENT = 6.0, GRACE_STEPS = 0, TAU = 0.5, IMPROVEMENT_EPS = 0.005, floors unchanged. Selection screened 30 of 178 unique candidate configurations; 29 were admissible at all four seeds (20260725, 7, 424242, 20270101). The audit's false-optimism instrument also gained an absolute bar of 0.25 on delivered credit, independent of max_recovery_credit, closing the measurement defect whereby the bar scaled with the constant under test.

Audit outcome at the shipped constants (seed 20260725): R_f = 0.7769 (band [0.731, 0.931]), false-optimism rate 0.0182 (absolute bar 0.25, rate bar 0.20), collapse fraction 0.0961, safety_net_required = false. The safety-net obligation ADR-031 left outstanding is discharged as measured by the audit lens.

Backward compatibility: breadth_exponent = 0.0 with no stagnation state reproduces v2.2 exactly. Profiles with every construct above its floor are bit-identical to v2.2 (§3.5 Case 1 is unchanged); only below-floor recovery moves.

Implementation and parity status: src/inference/welfare_scoring.py is now the source of truth for v3 semantics. spaces/maninagarden/welfare.py is synced to it and covered by an automated src↔spaces parity test. The HF Jobs training template inside spaces/maninagarden/training.py still carries an inline v2.2 copy of the formula and is not parity-tested; forecaster labels generated from it therefore encode v2.2 recovery credit and predate the gates — the same class of staleness this changelog notes for ≤v2.1.1 checkpoints. Advancing that template is a recorded successor obligation, as is threading stagnation counters through the live investigation loop (production call paths currently receive breadth gating only).

Decision record: ADR-032. Audit and amendment: ADR-031.

Version 2.2 (2026-07-08): Bounded, One-Sided Penalties, Recovery Guards

Remediation release following a second adversarial review that executed the paper's §5.1 reference implementation against its own claims. Supersedes two v2.1.1 stances; retains and extends the rest.

  1. Normalization (§3.1): v2.1.1's Fix #1 documented Φ ∈ [0, ~1.48] but left the function unnormalized. v2.2 divides by Φ_max = 1 + 4η + η_curiosity = 1.48, so a perfect society scores exactly 1.0. Division by a positive constant preserves all orderings. Supersedes Fix #1's documented-but-unnormalized stance.

  2. One-sided primary penalties (§3.4): the symmetric squares punished love-rich/resource-poor communities for their poverty a second time. The four primary pairs now penalize only the theorized failure direction; the curiosity cross-pair stays two-sided (both its directions are named failures). The Pareto-efficiency claim inherited from the Nash SWF is retracted and replaced with tested guarantees (§1.2 caveat).

  3. Recovery guards (§3.1.2): dx/dt is clipped to ±0.3 and recovery credit capped at 0.5 — a collapsed construct's effective value can no longer be lifted to its floor by a single optimistic trajectory input (raw 0.01 + dx/dt=1.0 previously read as exactly at-floor). Fix #3's lagged λ_L (circularity breaking) is retained.

  4. Priorities, not gradients (§6.1, §3.3.1): v2.1.1's Fix #4 computed the true ∂Φ/∂x numerically and clamped it to ≥ 0 on the claim that Φ is monotone in each construct — which the penalty term makes false. The clamp returns zero priority precisely in divergence regions, deprioritizing investigation where structural distortion lives. v2.2 reverts to an explicit scarcity-weighted proxy (construct_priority = λ_L^0.5·wᵢ/xᵢ), honestly documented as not-the-gradient. The stale example values (1.43/0.16, from v1.0's exponent formula) are replaced with measured v2.2 values (2.95/0.06). Supersedes Fix #4.

  5. Fail-loud inputs (§5.1): compute_phi raises on missing constructs instead of silently defaulting them to 0.5 — a missing measurement is an information gap.

Contract locked in property tests: tests/inference/test_welfare_properties.py (bounds, monotone directions, recovery caps, fail-loud, paternalism ordering, src/spaces implementation parity).

Implementation: spaces/maninagarden/welfare.py, src/inference/welfare_scoring.py, and the HF Jobs training template all carried the v2.2 changes (including lam_L_prev). Forecaster checkpoints trained on ≤v2.1.1 outputs should be retrained before mixing with v2.2 data. (See v3.0 below for the current reference and for which copies are parity-tested — as of v3.0 the training template has not been advanced.)

Version 2.1.1 (2026-03-08): Mathematical Audit Fixes

Patch release: Ten fixes from systematic mathematical audit of the formula and documentation.

Code changes (src/inference/welfare_scoring.py):

  1. Lagged λ_L recovery (Fix #3): compute_phi() accepts optional lam_L_prev parameter. When λ_L is below its floor, its own recovery uses the lagged value λ_L(t-1) instead of the current value, breaking the circular self-reference. Backward-compatible: None defaults to current λ_L.
  2. Numerical gradient (Fix #4): phi_gradient_wrt() replaced analytical approximation (solidarity * w_i / x) with central finite differences (Φ(x+ε) − Φ(x−ε)) / 2ε (ε=10⁻⁵). Captures synergy, penalty, recovery floor, and equity weight redistribution effects that the analytical form missed. Gradient clamped to ≥ 0.

Documentation changes (this file): 3. Φ range corrected (Fix #1): Φ ∈ [0, 1.48], not [0,1]. Ψ_ubuntu ≥ 1.0 pushes maximum above 1. 4. Equity weight trade-off (Fix #2): Acknowledged partial substitutability as known trade-off (§3.1.3). 5. λ_L dominance documented (Fix #5): Quadruple influence through 4 channels documented as intentional Ubuntu design (§3.1.1). 6. Penalty bound proven (Fix #6): Formal proof that Ψ_penalty ≤ μ = 0.15, so (1 - Ψ_penalty) ≥ 0.85 (§3.4). 7. Case 4 fixed (Fix #7): Changed dx/dt from 0.3 to 0.0 to properly demonstrate community-mediated recovery (§3.5). 8. Double-counting justified (Fix #8): Synergy-penalty overlap on (c,λ_L) pair documented as intentional paternalism/white supremacy detection mechanism (§3.4). 9. Sensitivity analysis (Fix #9): Parameter sensitivity table added (§7.5). γ is the most sensitive parameter (3.4% per ±10%). 10. Static snapshot (Fix #10): Φ acknowledged as static function; temporal dynamics require PhiTrajectoryForecaster (§7.6).

Version 2.1 (2026-02-26): Recovery-Aware Floors + Equity Weights

Major revision: Three structural additions to the formula.

  1. Recovery-aware floors (§3.1.2): Below-floor constructs receive community-mediated recovery potential via trajectory (dx/dt) and community capacity (λ_L^0.5). Key insight: care doesn't begin the uptick without community intervention. Sigmoid bias of −3.0 prevents trajectory from dominating community capacity.

  2. Equity-adjusted weights on effective inputs (§3.1.3): Weights wᵢ = (1/x̃ᵢ)/Σ(1/x̃ⱼ) computed on recovery-adjusted values, not raw metrics. Supersedes the fixed per-construct exponents (α=0.7/0.8/1.0) from v1.0-v2.0. Achieves Rawlsian maximin dynamically.

  3. Curiosity cross-pair (§3.3.1): Love × Truth diagonal synergy with η_curiosity=0.08. Divergence penalty expanded from 4 to 5 pairs, adding (λ_L − ξ)² for surveillance/willful-ignorance detection.

Implementation: spaces/maninagarden/welfare.py was the reference implementation at this version (formula v2.1-recovery-floors). Training job launched with recovery floors active in data generation for the first time.

Philosophical grounding:

  • Recovery floors: Nussbaum (2000) non-negotiable capability thresholds + Ubuntu community-mediated recovery
  • Equity weights: Rawls (1971) maximin via dynamic inverse-deprivation
  • Curiosity: hooks (2000) love as extension for growth + Fricker (2007) epistemic integrity

Version 2.0 (2026-02-19): Love/Protection Split

Major revision: Split λ into λ_L (love) and λ_P (protection)

Rationale: Original conflated construct merged generative developmental support (hooks' love) with defensive safeguarding (protection)—philosophically distinct phenomena requiring separate measurement.

Theoretical grounding:

  • hooks (2000): Love as active extension for growth
  • Berlin (1969): Positive vs. negative freedom
  • Paternalism detection: high c+λ_P, low λ_L

Structural changes:

  • Constructs: 7 → 8
  • Nash weights: 1/7 → 1/8
  • Synergy pairs: c×ε → c×λ_L, κ×λ → κ×λ_P
  • Worked examples recalculated

Mathematical verification: All gradient properties preserved. Nash SWF structure intact.

Version 1.0 (2026-02-18): Initial Formalization

  • Original 7-construct model
  • Nash SWF structure
  • Atkinson inequality sensitivity
  • Synergy/penalty terms

Acknowledgments

This work builds on decades of scholarship in welfare economics (Atkinson, Sen, Rawls), feminist epistemology (Collins, hooks, Fricker), and AI alignment (Russell, Anthropic). Special thanks to bell hooks, whose definition of love as active extension opened the path to formalizing care ethics mathematically.

Funding: Independent research, no external funding.

Conflicts of Interest: None declared.

Data Availability: Reference implementation available at https://github.com/crichalchemist/wave-experiment


For correspondence: [Contact information]

Suggested Citation:

Φ(humanity): A Rigorous Ethical-Affective Objective Function. Working Paper v3.0. (2026). Detective LLM Project.


This function is not a turnkey moral oracle. It is a disciplined framework forcing transparency about normative commitments while structurally preventing care-without-love dystopias and centering marginalized voices in the definition of flourishing.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train crichalchemist/phi-humanity-welfare-function

Papers for crichalchemist/phi-humanity-welfare-function