Project SHAME - Sustainable Human Accountability Metrics Engine

Sorry for the interlude… It’s my writing style…

Next up… MORALS! ^^

Now this was a long time ago. I put it on my MP3 player with TTS and walked round the shire considering it. :confused:

(Of course others thought all this before me… I’m just trying to explain this as a project to my children and posting it publicly like this makes it ‘dual use technology’.)

Now it really wouldn’t be right to create an PhPU for the kids for Morals…

Moral reasoning is not something you offload to a machine or to a business…

In fact, as we discussed… Different people may read the same story very differently depending on where they come from…

For example… In this story my mum was confused that the story was about a man when it was clearly, to her mind, a girl…

To the kids, brought up with my wife watching Chinese dramas, they were clear this was not the case.

Now this gives us an intensely interesting perspective - and project…

How do different AI reasoning models interpret the same stories… and why?

Enter… the major conflict in my children’s lives… their dual English / Chinese heritage…

Teaching them “intelligently” does not mean asking AI to decide the moral of a story. It means comparing its answers, uncovering its assumptions, and recognising the limits of scaling moral reasoning through machines and businesses…

And furthermore… questioning what happens when entire collections of “moral stories” are written through the perspective of a single AI model.

The Moral Scaling Problem

So far, we have been producing our moral stories one at a time:

  • Generating the story

  • Creating the images

  • Checking the audio

  • Writing the YouTube description

Yesterday, my dad asked me to automate the process.

We can now generate up to 20 stories in a single batch.

The process is now fairly simple: select the stories we want, click Submit, and leave the system running while we work on something else. A full batch currently takes approximately three hours.

The pipeline uses:

  • GPT OSS 120B: Fetches story titles and writes the stories.

  • Qwen Forced Aligner: Matches each word with its correct timing in the audio.

  • Qwen TTS Peter Clone: Converts all 200 generated audio files to use a voice based on my dad’s.

Things still to fix

  • The DGX Spark currently loads and unloads the models inefficiently. (Unloads and loads the same models between stories).

The important thing to note is that we still have to review these stories.

  • So another problem is the necessity of human reviewing -
  • And just because we reviewed it doesn’t mean we reviewed it right.

“So another problem is the necessity of human reviewing -And just because we reviewed it doesn’t mean we reviewed it right.”


You could do a version of falsification. Ask how could the moral be wrong or where could it be misapplied. I do a lot with constraints.

Out of hearts but I’ll be back :infinity::heart:

This is how I use it, you can do it the same way just applied to morals. Just ask a google it can direct you to the methodology and explain it. All my stuff is open source so please use what you want or don’t :winking_face_with_tongue:

We already have a small system for comparing answers across models, and I agree that falsification could strengthen it by asking where a moral might fail, be misapplied, or produce unintended harm. These questions are already partly included in the Moral Compare tab shown in the attached image.

However, I think moral interpretation presents several deeper problems:

1) The composition, omissions and weighting of model training data are not fully visible, whether we are examining an internal model or an external one. This may affect interpretation in many ways; for example, a model developed within one society may interpret a story differently when it is applied within another, particularly across Eastern and Western cultural traditions.

2) A model’s moral interpretation is not equivalent to a person’s moral judgement. Local, cultural and minority perspectives may be underrepresented or averaged away.

3) Model outputs are non-deterministic generated results. Unlike an individual, whose views may have developed through relatively stable life experiences, a model can produce different—and sometimes opposing—interpretations from similar prompts.

4) The context in which a story is read may also shape the reader’s interpretation. Age, current social dynamics, personal experience and what someone recognises in the situation at a particular point in their life can all affect the moral they take from it.

We can therefore compare models, test assumptions, search for counterexamples, examine cultural context and ask where each moral might be misapplied.

But falsification does not necessarily reveal one final “correct” moral. It can help us identify interpretations that are weaker, narrower or potentially more harmful, while still leaving room for legitimate differences in human perspective.

That is really the purpose of the system: not to make the machine the moral authority, but to make its assumptions visible enough for people, particularly children and parents, to examine and discuss them.

It’s like art man, in the ear of the hearer ie eye of beholder.
Asking how it can be wrong is more solid than meaning because what it means to one person is 100% different than someone else despite language. Then once you toss language on top of it. Universal things like food must be ate . Air must be breathed how do they fail when one commands no one gets food or when air is un breathable

Hunger, cold, wet, hurt, dead are all universal to the human condition

You as the parent should set your morals and then they would be reflected in the machines you build anyway..

A moral I live by .. the world is a harmful place I try to not add unnecessary harm.

Survival and evolution are never clean is a good one too.

Model is less than self is good.

And everything we think is real is a model of a mind modeling reality

Trying not to add unnecessary harm is probably one of the firmer shared principles we can begin with, even though people will still disagree about what harm is necessary or avoidable.

I also agree that parents inevitably influence both their children and the systems they build. But I think parents must first accept that their own beliefs must remain open to falsification too.

Our moral views are partial, shaped by culture, experience, institutions and the world in which we grew up. Some of what we believe may be wrong, incomplete or unsuitable for the world our children will inhabit.

We should certainly explain our beliefs, values and reasons to our children. But encoding our worldview into machines and then using those machines to reinforce it risks automating inherited bias rather than teaching moral judgement.

Our children will not live in the same world we did. The aim is not to leave them without guidance, but to give them the tools to question, compare and eventually take ownership of their own values.

Guidance without the freedom to examine it can easily become indoctrination.

I love your project Phyde. I’m a fan

It is a needed area that is understudied family and AI dynamics I always felt you were suited to this kind of work.

You really should use Medium too Phyde, this would be a huge hit on it. It is free

Really fun idea! it is important to educate kids about how fundamentally different machines really are.

A meme I think is funny and true.

From individual humans, yes. From systems, not necessarily.

That meme captures a much wider problem. People often treat a system’s confident answer as evidence that understanding must exist somewhere within it.

Systems reduce cognitive load. They package judgement into instructions, headlines, scores, procedures and apparently simple answers. That can be extremely useful, but it can also encourage us to accept the output without examining the machinery beneath it.

Modern compulsory mass education is itself relatively recent. One of its continuing tensions is whether education teaches people to question systems, or primarily teaches them how to function successfully within them.

Maybe, I think these systems can end up making us less human.

What is “base” human to be lesser than?

Just thinking along that path… What is “base human” we must ask..
What qualities define being fully human?
Who decides that those qualities are the standard?
If technology changes how we think and act, is that becoming “less human,” or just becoming a different kind of human?

Perhaps the better question is not what is a base human?

It is when do humans begin thinking like the systems they inhabit, rather than questioning them?

We think like the system because the system has constraints real rules one can not break…

I’m out of hearts I appreciate the conversation @phyde1001

We question systems, but we also have to recognize which constraints are fundamental and which are merely conventions. :infinity::heart:

Sorry MD tab is easy to hit on my phone but hard to hit when I want to…

You guys are great to read! I agree with Mitchell if you go on medium you would be read Phyde, my name is quad on it. I would “read list” your writing for my readers. Mitchell is Mitchell everywhere :laughing:

Thank you, PandaPi. I appreciate that.

Unfortunately, I am currently navigating some fairly fundamental constraints that I hope are temporary.

I can see in real life how quickly children absorb the assumptions of the systems around them. Like a plant bending towards the light, sometimes the pot must be turned—not to prevent growth, but to ensure that it does not occur in only one direction.

We are hoping to return to China at the end of the month. I may not initially be able to work there because of points and degree requirements, but I hope to continue writing and developing my personal projects.

Medium is not a practical option for me in China, so I hope this Discourse forum remains accessible and I can continue contributing here through projects such as SPARK and Agent GIF.

For me, it is less about being widely read than about producing work that others can examine, verify, use and build upon. Even if no one else does, my children will have a public record of what I tried to build and why.

One of my greatest fears about AI is that we may see a fractal pattern resembling the emergence of modern compulsory mass education.

In the UK, compulsory schooling for the masses was established only a few decades before the First World War. I am not suggesting that one caused the other, but within a generation, societies had developed extraordinary new capacities to educate, organise, persuade and mobilise populations at scale.

The same systems that can expand human capability can also make populations easier to coordinate and direct.

That is the fractal I worry about with AI: an extraordinarily powerful force for good whose structure may also scale conformity, persuasion and control.

Rather SHAMEful, if we fail to account for that possibility.

My son has just completed an intent router as part of this project but I have encouraged him to put it under his own thread.

When worlds collide, and multilingual families are separated, how do you protect your children’s futures, keep families interacting, and prevent the problems of the adult world from tearing the child apart?

Children should not have to lose half of themselves because of adult failings—especially when so much has already been fought for between two worlds.

When language is not fluent between those worlds, how can one parent help fill the gap?

Ultimately, sometimes one side has to take the hit so the children do not. But perhaps AI can soften that transition.

This is an early experimental method that may help other families fare better.

I do not believe AI can eliminate the distance between cultures, parents or families. But even where the immediate family is stable and close, the wider family may still be thousands of miles away. Perhaps tools like these can help parents protect a child’s language and maintain a bridge to the other half of their world.

I can check and verify the English reasonably well. What I cannot yet do is independently verify the Chinese with the same degree of confidence.

So this is our current AI bridging method for checking the Chinese output:

  1. Generate the Chinese audio using your preferred TTS model.
    In this case I am using Qwen3-TTS with a custom voice, partly to give the model a proper test.
  2. Feed the generated audio into Whisper.
    Whisper independently transcribes what it hears back into Chinese text.
  3. Compare the transcription against the original Chinese text.
    Any missing, substituted or substantially different characters can then be flagged for further checking.

This is not full verification.

If both systems make compatible mistakes, agreement between them does not prove that the Chinese is correct.

More importantly, this round-trip test checks whether the spoken Chinese remains faithful to the Chinese text. It does not prove that the original Chinese translation itself is semantically correct.

What it does give us is a second, independent signal that should catch at least some pronunciation errors, omissions and unintended changes that would otherwise pass unnoticed.

The process could be strengthened further by using multiple speech-recognition models, multiple TTS models, or—most importantly—a human Chinese verification step.

Until now, our process has effectively been:

Chinese text → Chinese TTS → assume the spoken Chinese is correct.

Now it becomes:

Chinese text → TTS → audio → Whisper → recovered Chinese → comparison with original.

That still does not give us certainty.

But it gives us something much more useful than assumption:

an automated second opinion.

Here is a live example from The Boy Who Cried Wolf.

The original Chinese text was:

从前有一个放羊的男孩,他负责看守羊群

The generated audio was then passed back through Whisper.

Whisper recovered:

从前有一个放羊的男孩,他负责看守羊群

Result: exact textual match, apart from punctuation.

That does not prove the translation itself is perfect, nor does it prove every aspect of pronunciation is correct.

But it does demonstrate that the generated speech survived the round trip through an independent speech-recognition model without changing the underlying sentence.

Up until this point, we were effectively using the Chinese TTS model while assuming its Chinese output was at least as reliable as its English.

Adding Whisper gives us a way to automatically challenge that assumption.

It is still imperfect.

It is still experimental.

And ultimately, human verification remains the strongest check.

But when you are trying to preserve a child’s connection to both sides of their family, even an imperfect bridge may be worth building.