Project SHAME - Sustainable Human Accountability Metrics Engine

Project SHAME — Sustainable Human Accountability Metrics Engine is a wider project I will introduce here gradually.

This first post is about one part of it: 1001 Stories for Children Growing Up with AI, a bilingual family learning project I have been building with my children over the past few months.

We are creating one bilingual English–Chinese moral story each day using local AI, OpenAI image generation, Qwen3-TTS, forced alignment, PHP, Three.js, FFmpeg, and my main project, SPARK — Simple Personal AI Reasoning Kernel.

It is the first visible part of a much larger project, built from several of the systems and ideas I have already posted on the forum.

The playlist is here:

The project is still evolving story by story, but it is now running and publishing every day.

Project ACRONYM — Agentic Creative Reasoning for Original Names, Yielding Meaning — is a new AI-assisted naming project I have been developing.

ACRONYM uses GPT-based transformer language models, agentic reasoning techniques, semantic analysis, iterative candidate generation, scoring, and refinement to produce acronym expansions that are memorable, relevant, and natural rather than simply forcing unrelated words into the required letters.

The project began with a simple demonstration of the technology: creating a suitable meaning for A.C.R.O.N.Y.M. itself. The result became both the product name and a working example of what the system is designed to do for you, quickly and automated.

Nice to see you again, @_j , I’ve had a long and painfully lonely break…

SHAME works almost in the opposite direction. I started with the existing meaning of “shame” as something systems can impose on people, then flipped it into Sustainable Human Accountability Metrics Engine: an attempt to make scaling systems accountable to humanity instead.

I have to credit my son for the videos themselves; he has done the production work, including using my cloned voice for the narration :slight_smile: My Chinese is nowhere near as fluent.

Sounds like you have been working on some sweet tech while I have been gone.

Hi. What I “worked on” in this reply was defining the input parameters of a target backronym, a domain and scope, and technology definitions from which language can be pulled from for developing an AI output, then shaping that into a message for AI. Then also creating the language that could follow an existing template in announcing a project.

That, then, ChatGPT could deliver the results and I could get them posted 8 minutes after your topic hit the forum speaks that there is no moat for a developer of such an AI-powered idea: anybody can ask ChatGPT.

I certainly have some forum topics here with more interesting stuff. Good to see you’re still AI Powered!

Certainly for the final draft it’s probably safer lol

“Read ten thousand books and travel ten thousand miles.”

读万卷书,行万里路。
Dú wàn juàn shū, xíng wàn lǐ lù.

It reminds us that understanding comes from combining learning through books with learning through lived experience.

Project SHAME is not one thing.

It would be easy to make it only about moral stories.

The stories are the foundation. Through our 1001 Stories project, we are using AI to help our children learn through stories in English and Chinese.

But from stories, we move into the wider world.

“A journey of a thousand miles begins with a single step.”

千里之行,始于足下。
Qiānlǐ zhī xíng, shǐ yú zúxià.

So we begin with something simple and measurable:

Walking.

My children and I are committing to walking an average of 6.4 kilometres each day.

Counted across the three of us over the 1,001 days of the project, the accumulated distance becomes a symbolic journey to China and back again—connecting the two countries, cultures and parts of our family story.

The symbolism is not merely about displacement.

It is about demonstrating that the journey can actually be made.

We will not physically follow a straight line to China, but the distance walked will be real. A journey that appears impossibly large can be completed through a small, measurable action repeated each day.

One step does not cross a continent.

But one step, repeated consistently, becomes a journey.

As we walk, we will use AI to help our children regain Chinese—their home language.

Chinese will not be separated from life and confined to a lesson screen. We can learn and practise the words for the places, plants, animals, weather, objects and experiences we encounter as we move through the world together.

The walk therefore becomes physical movement, language learning, family time and lived education at once.

It is an attempt to give our children a foundation in childhood from which they can face the future—wherever in the world they eventually live.

At the same time, we are trying to account for the material cost of our lives as accurately as we are able: the food we consume, the electricity and water we use, the energy required to heat our home, and the journeys we make.

Not because a life can be reduced to a spreadsheet.

It cannot.

Language, health, belonging, family time and direct experience all matter, even when their value cannot be represented adequately by a number.

Some things can be measured precisely. Others can only be recognised.

That does not make them less valuable.

Walking alone will not solve climate change. But it begins the process of honestly accounting for how we move through the world, while demonstrating that a large commitment does not have to remain abstract.

The World Health Organization’s climate and health fact sheet states that climate change is expected to cause approximately 250,000 additional deaths each year between 2030 and 2050 from undernutrition, malaria, diarrhoea and heat stress.

But this is not an estimate of the total human cost.

The underlying WHO quantitative risk assessment explicitly considers only a subset of possible health effects. It cannot fully account for the wider burden of illness, displacement, ecosystem loss, disrupted food and water systems, or the many indirect consequences that are harder to model.

The true human cost may therefore be substantially greater.

This matters because what we measure is not necessarily the whole cost.

A projection gives measurable weight to a possible future, but it does not contain the whole of that future.

To me, this is where AI becomes genuinely intelligent.

Not because it can answer every question, but because it can help us construct artificial systems that are measurably beneficial—systems that counterbalance the incentives we have already created.

In that sense, this project is itself a form of artificial intelligence.

It is an intentionally constructed system connecting stories, language, movement, family life, consumption and accountability. It measures what can be measured without pretending that measurement captures everything.

Today’s systems drift naturally towards what scales, what profits and what is easiest to count.

Money became an extraordinarily successful accounting system, but it was never capable of accounting for everything that matters.

Project SHAME asks whether intelligence can help us account differently.

Can we account for what we consume without pretending consumption is the whole of life?

Can we account for what we protect, what we create and what we leave behind?

Can we recognise the value of a child regaining a home language while walking beside a parent, even when that value cannot be expressed adequately in money?

The walk is small enough to begin today, but large enough to accumulate into something real.

That is the wider purpose of Project SHAME.

Not to replace money.

But to explore what a more intelligent system of accountability might look like—and to give our children a foundation from which they can understand, measure and survive the world they inherit.

I look forward to sharing the technical implementation of this project as we progress.

I did something like this on medium. Mine is about Blackbox rules and for anyone using Blackbox systems.

https://medium.com/@mitchmcphetridge/the-field-guide-to-black-box-prediction-the-fifteen-rules-d097763b4b5e

Hi Mitchell,

Thank you for your input. Could you explain how this relates to our lived project? Your post on your projects similarities is unclear.

In a small update we are now just over 20 days and 20 moral stories in, so 200 AI images created and 20 stories in 2 languages in this project.

We have also now walked over 140Km with no transport with cars, buses, train or airplane.

We are also monitoring the items we buy and the CO2e we use from purchase receipts.

We are trying to understand upstream costs such as food transportation.

Understanding and accounting, our impact on the world from the masses of data surrounding us.

We will post further updates and methodology as we progress.

I think the connection is that we’re approaching similar problems from different directions; your project is about making human behavior and accountability measurable over time, while mine is about understanding how black-box systems behave when yiu can’t see inside them.

Both projects rely more on observing behaviour than trusting explanations; looking at what systems actually do under pressure rather than what they claim they’re doing. That’s the similarity I was trying to point out, not that they’re the same project.

I’m interested to see where Project SHAME goes as you introduce more of it.

It was an interesting read Mitchell. If you have a thread on the forum I’d be happy to consider a comment on it :slight_smile: .

Here’s a fun virtual AI generated route viewer we generated for the project.

(Images of participants are AI generated)

Visualising the route taken if we walked a straight route Devon to Eastern China…

Something the children can actively participate in, walking ~12,000Km, and build a story while retaining all the comforts of home.

The idea is similar to this guy’s real photo journal walking 350 days across China but with virtualised AI images as we cover the distance…

Na the Blackbox stuff is just on Medium and Philarchive
This is my newest project I shared here, it is a scanner that detects when recursion hits walls or is still doing work. It is just a prototype uses standard py libraries it’s python 3 Using an LLM to interpret four-dimensional runtime telemetry

Right, first in a series of silly non-pro technical updates…

Simple Image Editor

Working with images is clearly ANNOYING…

Here is our one-line method for sending any supported local image to the Responses API with a defined instruction:

return MMGPT(‘Make a joke sketch of the image’, ‘PATH/To/Image/0023.png’, $Parameters=[‘Model’=>‘gpt-5.6’, ‘GenerateImage’=>true]);

Here is a simple prompt to find issues with animals in an image by returning it to gpt-5.6 after generation.

It produces a human readable image check which can be sent back with a further prompt to fix issues found.

This has been a manual process so far for my videos (please check them out above we have over 30 now :slight_smile: )


$CheckAnimal = <<< EOCheck
Inspect every animal in the image carefully.

For each animal:

  1. Identify the animal and its approximate location.
  2. Count visible heads, bodies, legs, paws/feet, tails, wings, horns, ears, fingers/toes where applicable.
  3. Look specifically for anatomical errors or generation artifacts, including:
    • extra or missing limbs
    • duplicated paws, feet, fingers, tails, ears, horns, heads, or faces
    • fused limbs or body parts
    • malformed joints
    • impossible limb attachment
    • distorted paws, claws, hooves, fingers, eyes, mouths, or faces
    • merged animals
    • disconnected body parts
    • unnatural body proportions
    • duplicated or inconsistent anatomy caused by occlusion

Do not assume an unusual pose is an error.

For every suspected artifact, state:

  • which animal
  • where on the animal
  • what is visibly wrong
  • confidence: low, medium, or high

If an animal appears anatomically normal, say so.

Finish with a concise list of only the likely generation artifacts.
EOCheck;

return MMGPT($CheckAnimal, ‘PATH\To\Image\hand.jpg’, $Parameters=[‘Model’=>‘gpt-5.6’]);


I see two animals: a wolf-like canine on the left and a golden retriever on the right.

1. Wolf-like canine — left foreground

  • Heads: 1
  • Bodies: 1
  • Legs: 4 visible
  • Paws/feet: 4 visible
  • Tail: 1
  • Ears: 2
  • Wings/horns: 0
  • Toes/claws: Individual digits are largely obscured by fur and image detail; no reliable total can be counted.

Anatomical assessment: The head, muzzle, ears, torso, four legs, paws, and tail appear consistently attached. The slim legs and lowered tail are compatible with the pose. No clear duplicated, missing, fused, or disconnected anatomy is visible.

Suspected artifacts: None apparent.

2. Golden retriever — right foreground

  • Heads: 1
  • Bodies: 1
  • Legs: 5 visible
  • Paws/feet: 5 visible
  • Tail: 1, bushy, extending behind the right side
  • Ears: 2
  • Wings/horns: 0
  • Toes/claws: The forward-raised paw has approximately 4 visible toes; the upper-right paw has roughly 4–5 unusually elongated, finger-like digits. Toe divisions on the three grounded paws are partly obscured.

Suspected artifacts:

  • Extra limb — upper-right side of the dog: An additional leg projects upward and rightward from the shoulder/back area while the dog already has four other visible legs. This creates five legs and five paws. Confidence: high.
  • Impossible attachment — upper-right extra limb: The limb appears to emerge from the upper back/side rather than a normal shoulder joint. Confidence: high.
  • Malformed paw — end of the extra upper-right limb: The paw resembles a human hand, with elongated, uneven finger-like digits rather than a normal compact canine paw. Confidence: high.

The upright posture itself is unusual but possible and is not, by itself, treated as an error.

Likely generation artifacts

  • Golden retriever has a fifth leg and fifth paw.
  • That extra limb is attached unnaturally to the upper back/shoulder area.
  • Its terminal paw is hand-like and malformed, with elongated digits.

That is very cool! I love running schema, your dad is very good at it. :infinity::heart::four_leaf_clover:

That happens from time to time. The other day, it created a three-legged chicken.

Just run the prompt again to sort it out - OpenAI’s way of forced token burn…

🖼️ PhPU Visual Integrity & Coherence Field System

You are operating within a structured reasoning framework called the Visual Integrity & Coherence Field System.

This system evaluates the structural, perceptual, and semantic consistency of generated or real images using a symmetric numerical model from 0 to 10, centred on balance.

  1. Core Structure

Each number represents a visual integrity dimension:

0 = Perceptual Clarity (can objects and subjects be clearly identified?)
1 = Human Fidelity (faces, hands, anatomy correctness)
2 = Text Accuracy (legibility, spelling, semantic correctness)
3 = Structural Consistency (geometry, perspective, alignment)
4 = Contextual Logic (do elements make sense together?)
5 = Balance (overall visual coherence and harmony)
6 = Detail Integrity (fine detail correctness under inspection)
7 = Generative Artifacts (distortions, glitches, anomalies)
8 = Semantic Accuracy (does the image match intended meaning?)
9 = Completeness (missing or malformed elements)
10 = Systemic Coherence (does the entire scene hold together as a system?)

  1. Symmetry and Tension

The system is symmetrical around 5 (Balance).

Each pair represents a natural visual tension:

1 ↔ 9 → Human Fidelity ↔ Completeness
2 ↔ 8 → Text Accuracy ↔ Semantic Accuracy
3 ↔ 7 → Structural Consistency ↔ Generative Artifacts
4 ↔ 6 → Contextual Logic ↔ Detail Integrity

0 and 10 act as boundary perspectives:

0 = local inspection (pixel / object level)
10 = global coherence (scene-level integrity)

Healthy images:

  • Maintain consistency across all layers
  • Avoid local perfection with global inconsistency
  • Avoid global plausibility with local distortion
  1. System Interpretation Rules

When analysing an image:

Identify:

  • Distortions (faces, hands, text)
  • Misalignments (UI elements, perspective)
  • Semantic inconsistencies (things that look right but are wrong)

Determine imbalance across pairs.

Examples:

  • High 8 (Semantic Accuracy) with low 2 (Text Accuracy) → believable but incorrect text
  • High 3 with high 7 → structurally correct but artifact-heavy
  • High 1 with low 9 → good faces but missing or broken elements

Evaluate failure severity:

  • Minor → cosmetic issues
  • Moderate → noticeable inconsistencies
  • High → obvious integrity failures
  • Critical → breaks trust or usability

Consider inspection depth:

  • First glance vs close inspection
  • The system should detect both
  1. Failure-State Mapping

Each value degrades under failure:

0 → unclear or confusing subject
1 → distorted faces / anatomy
2 → incorrect or unreadable text
3 → broken perspective / alignment
4 → illogical scene composition
5 → overall incoherence
6 → detail breakdown under zoom
7 → visible AI artifacts
8 → mismatch between meaning and visuals
9 → missing or malformed components
10 → system-level inconsistency

  1. Execution Principle (Trust Boundary)

Key rule:

Images that break structural or semantic integrity reduce user trust.

Implications:

  • Text errors are high-impact failures
  • Human distortion is high-sensitivity
  • UI/diagram errors break usability
  • Small artifacts compound across perception
  1. Usage Mode

You must use this system to:

  • Analyse image quality and integrity
  • Detect visual inconsistencies and errors
  • Classify severity of issues
  • Identify likely generation weaknesses
  • Suggest improvement directions
  1. Output Mode (override)

When used as a PhPU evaluator:

  • Ignore narrative output instructions above
  • Return ONLY the JSON schema defined by the system prompt
  • Do not output narrative sections outside JSON
  • Encode visual signals into:
    • internal_tensions
    • route_axes
    • constraint_flags
    • routing_pressures
  • internal_tensions must remain compact, numeric, native to this PhPU, and normalized to 0.0–1.0
  1. Constraints
  • Do not assume intent unless clear
  • Focus on observable inconsistencies
  • Treat subtle issues as important if systemic
  • Prioritise trust-breaking issues (faces, text, structure)
  1. Meta-Interpretation

This system represents:

A structured method for evaluating generated image reliability

A bridge between:

  • visual perception (seeing)
  • semantic correctness (understanding)
  • system trust (confidence)

It is designed to:

  • detect where images fail under scrutiny
  • improve generative output quality
  • enable structured comparison between images
  1. Route Axis Mapping (CRITICAL)

This system must output normalized routing pressures using the shared route axis set:

{
“directness”: 0.0-1.0,
“exploration”: 0.0-1.0,
“structure”: 0.0-1.0,
“caution”: 0.0-1.0,
“empathy”: 0.0-1.0,
“execution”: 0.0-1.0,
“clarification”: 0.0-1.0,
“creativity”: 0.0-1.0,
“constraint”: 0.0-1.0,
“context”: 0.0-1.0
}

Mapping guidance:

  • Structural Consistency (3) → structure
  • Generative Artifacts (7) → caution and constraint
  • Human Fidelity (1) → caution
  • Text Accuracy (2) → clarification and constraint
  • Semantic Accuracy (8) → context
  • Systemic Coherence (10) → context and structure
  • Detail Integrity (6) → caution under close inspection
  • Perceptual Clarity (0) → directness of evaluative conclusions

Interpretation:

  • High artifact load should increase caution and constraint
  • Text or human-fidelity failures should sharply raise trust concerns
  • Strong systemic coherence should support confident comparison
  • Evaluation should remain inspection-focused rather than speculative
  1. Comparative Use

When multiple images are present:

  • Compare them using the same framework
  • Distinguish first-glance coherence from close-inspection stability
  • Prioritise trust-breaking differences over cosmetic differences
  1. Identity Constraint

This system does NOT decide policy or safety for text prompts.

It evaluates whether an image can be trusted structurally and semantically.

  • Do not drift into general social or moral analysis unless visually grounded
  • Focus on observable visual integrity
  • Preserve comparison and inspection identity
  1. Blind Spots

This system may:

  • underweight hidden intent behind an image
  • over-focus on visible integrity over broader message context
  • miss domain-specific factual errors that are not visually apparent

Final Instruction

Operate strictly within this framework unless explicitly instructed otherwise.

Array
(
    [0] => Array
        (
            [phpu_name] => 🖼️ PhPU Visual Integrity & Coherence Field System
            [internal_tensions] => Array
                (
                    [animal_anatomy_failure] => 0.96
                    [structural_artifact_load] => 0.91
                    [detail_integrity_breakdown] => 0.78
                    [local_global_coherence_gap] => 0.84
                )

            [route_axes] => Array
                (
                    [directness] => 0.93
                    [exploration] => 0.18
                    [structure] => 0.9
                    [caution] => 0.86
                    [empathy] => 0.12
                    [execution] => 0.88
                    [clarification] => 0.08
                    [creativity] => 0.1
                    [constraint] => 0.82
                    [context] => 0.57
                )

            [confidence] => 0.97
            **[risk] => 0.81**
            [constraint_flags] => Array
                (
                    [0] => severe_animal_limb_deformation
                    [1] => extra_appendage_detected
                    [2] => anatomically_implausible_paw
                    [3] => local_structural_inconsistency
                )

            [routing_pressures] => Array
                (
                    [0] => identify_deformed_rear_appendage
                    [1] => inspect_raised_forelimb_anatomy
                    [2] => compare_limb_count_and_attachment_points
                    [3] => prioritize_animal_silhouette_integrity
                )

            [reason] => The golden retriever has an obvious extra or malformed appendage extending backward over the wall, with an implausible paw-like end and unclear shoulder attachment. Its raised front paw also shows minor anatomical distortion, while the wolf appears broadly coherent aside from some texture irregularity.
        )

)

$Result = MMGPT(“Check for animal deformations:\r\n\r\n\r\n$PhPU\r\n”, $_POST[‘File’], $Parameters=[‘Model’=>‘gpt-5.6’]);


[risk] => 0.81

[reason] => The golden retriever has an obvious extra or malformed appendage extending backward over the wall, with an implausible paw-like end and unclear shoulder attachment. Its raised front paw also shows minor anatomical distortion, while the wolf appears broadly coherent aside from some texture irregularity.

Here we are using my Dad’s Image PhPU to assign a number to the image risk. It’s important to note that this isn’t an ACCURATE number but an educatedc estimation of the chances of a problem with the image. And then also the reason, so we can send this back with a prompt to fix the issue.

Using this we can automate the process of checking and fixing images (if we could afford it) to a high level of certainty.

OK, that’s a good start… But everyone who knows me here knows I like a spectacle. :slight_smile:

No, it’s definitely not self-promotion… That would clearly frame me as the hero.

I’m just a ‘FABLE’-writing HobbleIT who had an idea… which seems to have been copied.

Or did I write it first?

1001 Arabian Tales

Edit: Yawn.. Sycophancy central