|
Proposal: Specialist Profiles / Execution Contracts for More Reliable ChatGPT Workflows
|
|
1
|
40
|
September 8, 2026
|
|
What makes humor a useful test case for language models?
|
|
0
|
72
|
July 13, 2026
|
|
Deprecation notice: Evals will be shut down on November 30th, 2026
|
|
6
|
436
|
July 3, 2026
|
|
OpenAI Evals Bug Unknown Parameter
|
|
3
|
222
|
May 22, 2026
|
|
OpenAI Platform Evals w/ gpt-5 models & file_search tool returns "Empty assistant message"
|
|
1
|
120
|
March 19, 2026
|
|
Evals Created Via API Cannot Be Opened in Dashboard UI
|
|
0
|
55
|
February 3, 2026
|
|
From GEO to DED: Governing How AI Represents Brands
|
|
0
|
325
|
February 2, 2026
|
|
Evaluations UI in the Dashboard are failing
|
|
11
|
420
|
January 29, 2026
|
|
When model answers are correct but still feel incomplete
|
|
6
|
181
|
January 21, 2026
|
|
Correct, Coherent and Still Incomplete: A Note on Context Prioritization in Language Models
|
|
0
|
86
|
January 19, 2026
|
|
Evals with custom endpoint model
|
|
0
|
94
|
January 14, 2026
|
|
OpenAI eval API - text similarity grading
|
|
2
|
127
|
December 10, 2025
|
|
Missing scopes: api.evals.delete
|
|
0
|
61
|
November 19, 2025
|
|
LLM and Prompt Evaluation Frameworks
|
|
13
|
14688
|
November 18, 2025
|
|
Create eval run via Node SDK fails due to incorrect property name: max_completion_tokens (incorrect) vs max_completions_tokens (correct)
|
|
2
|
131
|
November 13, 2025
|
|
Agent Builder Evals Incomplete
|
|
0
|
76
|
November 9, 2025
|
|
[Playground/Evals] How can I generate an output when a tool is invoked?
|
|
0
|
89
|
October 22, 2025
|
|
Evals with the responses data source always regenerate outputs
|
|
0
|
59
|
October 20, 2025
|
|
Evals Datasource API - Scheduled Evaluations
|
|
0
|
59
|
October 9, 2025
|
|
Evaluations - fails when importing from Logs
|
|
0
|
97
|
October 8, 2025
|
|
Evals product in Playground - Announcement and feedback
|
|
7
|
905
|
October 1, 2025
|
|
Chat Prompt Evaluation Unable to Reference Prompt ID or Saved Prompt on platform
|
|
2
|
188
|
October 1, 2025
|
|
New in Evals: Full Audio Support
|
|
0
|
142
|
September 12, 2025
|
|
Cannot set verbosity for gpt-5 evals
|
|
0
|
166
|
August 26, 2025
|
|
Evals: Invalid 'reasoning_effort' for non-reasoning model: gpt-5-chat-latest
|
|
3
|
1806
|
August 18, 2025
|
|
Is error 500 caused by JSONL file ID formatting?
|
|
3
|
130
|
August 15, 2025
|
|
BUG: Stored Chat Completions not showing in Dashboard when sending type "image_url" messages
|
|
5
|
478
|
August 8, 2025
|
|
Evals framework UI features changed not able to download results
|
|
5
|
446
|
August 4, 2025
|
|
Accessing Eval Results in OpenAI Agent SDK Response
|
|
0
|
200
|
August 1, 2025
|
|
Provide bulk test data for published prompt eval in Dashboard UI
|
|
0
|
103
|
July 30, 2025
|