Hi everyone,
I’m exploring a question at the intersection of language models and humor:
What makes humor a useful test case for language intelligence?
A joke is not just text. It depends on timing, surprise, shared context, ambiguity, culture, tone, social risk, audience expectation, and what the model chooses not to say. When models fail at humor, the failure can reveal where they miss human context.
I’m curious how people here would approach evaluating humor with language models.
For example:
- Can humor be evaluated beyond “is this funny?”
- What would a good humor eval look like?
- How would you test audience-specific reactions?
- Are agents useful for simulating different audience perspectives?
- Where do models tend to fail most visibly: timing, context, tone, cultural assumptions, or safety constraints?
I’m asking because I’m organizing a small AI + humor project/community effort and want to learn how builders think about this problem.
Would love thoughts, examples, or projects people have seen in this area.