Feature Proposal: Conversational Timing as Context

Bringing “Pauses” into Text-Based AI

Text-based AI can understand what a user says, but it often cannot see another part of human conversation:

the time before the user says it.

In human conversation, the same words can carry different implications depending on timing.

For example, an AI might say:

“There may be another interpretation of that decision.”

A user immediately replies:

“I see.”

Or the user remains inactive for 40 seconds and then replies:

“I see.”

The text is identical. But in human conversation, we would not necessarily interpret those two responses in exactly the same way.

Proposal

With explicit user consent, conversational timing could be provided to the AI as a weak contextual signal.

One particularly interesting signal is:

the time between the AI response being displayed and the user beginning to type.

The system should not directly interpret absolute timing.

For example:

“30 seconds = hesitation”

would be an unreliable assumption.

Instead, timing could be compared with the individual user’s own conversational baseline.

The model might receive only coarse information such as:

  • within usual range
  • somewhat longer than usual
  • significantly longer than usual

The AI would know that the conversational rhythm changed, but not why.

The AI should not decide what the pause means

A longer pause could mean many things.

The user may have been thinking deeply, uncertain, emotionally affected, distracted, answering a phone call, doing household tasks, or simply away from the device.

Therefore:

Timing should not become emotion diagnosis.

It should remain one additional piece of context that can be considered alongside the actual message and conversation history.

If the meaning remains uncertain, the AI should preserve that uncertainty.

Difference from typing-behavior analysis

Research already exists on keystroke dynamics, including typing speed, pauses during typing, corrections, deletions, cognitive load, expertise estimation, and behavioral biometrics.

This proposal focuses particularly on something that occurs before typing behavior begins:

the period in which nothing is being typed yet.

In simplified form:

AI response displayed
→ pre-typing interval
→ typing begins
→ typing behavior
→ message sent

The pre-typing interval may itself contain useful conversational information.

The proposal also emphasizes within-person deviation, rather than assuming that the same response time has the same meaning for everyone.

Privacy by design

This should not require storing detailed raw keystroke data.

For example, timing could be processed locally and converted into coarse signals such as:

usual / longer than usual / significantly longer than usual

before being provided to the model.

Raw keystrokes, drafts, or detailed behavioral histories would not need to be transmitted.

The feature should be opt-in and easy to disable.

Avoid changing the behavior being observed

If the AI constantly says:

“You took 27 seconds to reply.”

users may become self-conscious about their response time.

That could change the natural conversational behavior the system is trying to understand.

Therefore, conversational timing should usually remain a subtle background signal, rather than something the AI explicitly mentions in every interaction.

Goal

The goal is not to make AI “read minds.”

It is to recover a piece of conversational information that text interfaces often discard:

time.

A text-based AI already sees the words.

Perhaps it should sometimes also be able to notice the silence before them.

Not to decide what the silence means — only to avoid treating it as if no information existed at all.

Thanks for the thoughtful proposal. We’ll share the idea for optional, coarse timing context with local processing and clear privacy controls, while keeping the meaning of a pause uncertain. We don’t have a timeline to share.