# Confidence score for prompt response

**URL:** <https://community.openai.com/t/confidence-score-for-prompt-response/132278>\
**Category:** API\
**Created:** [March 31, 2023, 6:10pm UTC](https://community.openai.com/t/confidence-score-for-prompt-response/132278 "2023-03-31T18:10:21Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![wmarch015](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/wmarch015/32/33574_2.png) [@wmarch015](https://community.openai.com/u/wmarch015)\
**Post date:** [March 31, 2023, 6:10pm UTC](https://community.openai.com/t/confidence-score-for-prompt-response/132278/1 "2023-03-31T18:10:22Z")

</div>

Hi all,

I am trying the prompt response, where I used the opeai.completion.creat() to get response, I set the temperature = 0 to get the most probable answer, but I am curious whether I can get the confidence score/probability of the response. Is there any way to get this number?

Thank you.

---

<div class="post-metadata">

**Author:** ![hernandezpaul](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/hernandezpaul/32/186276_2.png) [@hernandezpaul](https://community.openai.com/u/hernandezpaul)\
**Post date:** [August 28, 2023, 12:12pm UTC](https://community.openai.com/t/confidence-score-for-prompt-response/132278/2 "2023-08-28T12:12:34Z")

</div>

Hi @wmarch015 ,  
could you please share with me if you have any progress on this topic.  
BR.  
Paul

---

<div class="post-metadata">

**Author:** ![jwatte](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/jwatte/32/127705_2.png) [@jwatte](https://community.openai.com/u/jwatte)\
**Post date:** [August 28, 2023, 3:18pm UTC](https://community.openai.com/t/confidence-score-for-prompt-response/132278/3 "2023-08-28T15:18:32Z")

</div>

The model doesn’t have any concept of “the response.”  
The model predicts one token at a time. It may have some “confidence” (or “inverse perplexity”) of _each token_ as it’s being generated, but it has no concept or control of the overall semantics of all generated tokens.  
Inside the implementation, the code that runs the model calls the model to generate one token, and then iterates, until it determines the model is done (either by the model outputting the “done” token, or by the generated tokens matching a stop sequence.)

If you could get the “confidence” of each token, perhaps you could multiply that value together for all the tokens you generated, but I doubt that would be a very meaningful answer – mainly, because the model isn’t built to have any concept of “the totality of the answer.” As soon as one token has been generated, that token goes back into the “history/input” side of the model, and is treated as “given context” for the next token.

---

<div class="post-metadata">

**Author:** ![hernandezpaul](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/hernandezpaul/32/186276_2.png) [@hernandezpaul](https://community.openai.com/u/hernandezpaul)\
**Post date:** [August 28, 2023, 6:25pm UTC](https://community.openai.com/t/confidence-score-for-prompt-response/132278/4 "2023-08-28T18:25:38Z")

</div>

Hi @jwatte,  
Great answer, thanks.

Do you know how is this related to the logprobs parameter available in the deprecated completions Api?

---

<div class="post-metadata">

**Author:** ![jwatte](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/jwatte/32/127705_2.png) [@jwatte](https://community.openai.com/u/jwatte)\
**Post date:** [August 29, 2023, 3:51pm UTC](https://community.openai.com/t/confidence-score-for-prompt-response/132278/5 "2023-08-29T15:51:53Z")

</div>

I have not used that one. From the documentation it sounds like it returns the probability of “the token” which might be “the last token it returned in the inference” but it seems a little ambiguous what it is it returns.

If there was something that returned the probability of each token that was actually chosen, that’d allow you to calculate the overall confidence, but it doesn’t sound like that’s what it is.

Separately: the “probability” here is very likely the output of a softmax operator (I don’t know if this is documented, but it smells like one,) which is nicely numerically behaved and tends to drive the model towards “making a choice,” but it’s not an exact percentage value of what it “should” be; it’s just an allocation of the choices that it it actually made. The model may be very confident, in the wrong answer 🙂 And why “predicted token weight” should map to “exponent of normalized weight” to generate “probability” is … hand-wavy. Seems to work OK in practice, though!

---

<div class="post-metadata">

**Author:** ![jwatte](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/jwatte/32/127705_2.png) [@jwatte](https://community.openai.com/u/jwatte)\
**Post date:** [August 29, 2023, 4:09pm UTC](https://community.openai.com/t/confidence-score-for-prompt-response/132278/6 "2023-08-29T16:09:40Z")

</div>

Alright, I tried the actual API, and it does return the list of logprobs for each token in both the input and output.

Thus, you could use that information to calculate the “probability” of this particular sequence being chosen, and using that as some form of “confidence.” But, in general, I wouldn’t put too much stock in that value, because it will be poorly behaved – for longer outputs, the probabilities will multiply out to lower overall probability, and if the model picks one low-probability token (which could be something unimportant like “and” instead of “the”) the overall probability will multiply out to much lower.

Anyway – if what you want is “confidence in the overall answer” then you can’t really construct that from “random probability of each word fragment token.”  
This actually gives a pretty neat insight into why I think these models aren’t really “thinking,” too – they just predict, one token after the next, with no “overall” model of what they’re doing.

---

<div class="post-metadata">

**Author:** ![anon10827405](https://avatars.discourse-cdn.com/v4/letter/a/ed8c4c/32.png) [@anon10827405](https://community.openai.com/u/anon10827405)\
**Post date:** [August 29, 2023, 4:13pm UTC](https://community.openai.com/t/confidence-score-for-prompt-response/132278/7 "2023-08-29T16:13:37Z")

</div>

> [@jwatte](#):
>
> Separately: the “probability” here is very likely the output of a softmax operator (I don’t know if this is documented, but it smells like one,)

The logprobs is the value before any scaling is applied. It’s for this reason you will see something like:

![Screenshot from 2023-08-29 12-13-23](https://us1.discourse-cdn.com/openai1/original/3X/d/2/d26b54c5691d38978ca0f73f773ea61effe075c8.png)

(if that’s what you were referring to) (correct me if I’m wrong) (it’s been so long since I’ve used davinci via api but I _think_ it’s actually returned as a value and the playground calculates the probabilities)

---

<div class="post-metadata">

**Author:** ![jwatte](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/jwatte/32/127705_2.png) [@jwatte](https://community.openai.com/u/jwatte)\
**Post date:** [August 29, 2023, 10:46pm UTC](https://community.openai.com/t/confidence-score-for-prompt-response/132278/8 "2023-08-29T22:46:55Z")

</div>

> [@anon10827405](#):
>
> The logprobs is the value before any scaling is applied

Yes, that makes sense; I agree!

---

<div class="post-metadata">

**Author:** ![avermeir](https://avatars.discourse-cdn.com/v4/letter/a/0ea827/32.png) [@avermeir](https://community.openai.com/u/avermeir)\
**Post date:** [August 30, 2023, 12:30am UTC](https://community.openai.com/t/confidence-score-for-prompt-response/132278/9 "2023-08-30T00:30:25Z")

</div>

That said, when you have a large body of responses, you can reinject them into the LLM and ask him to evaluate them.  
Additionnaly, you can ask the LLM to review his own response, or probe internet for fact checking what he said himself and give a confidence score based on that.  
Its all proxies but depending on your context it can help.

---

<div class="post-metadata">

**Author:** ![\_j](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/_j/32/766292_2.png) [@\_j](https://community.openai.com/u/_j)\
**Post date:** [August 30, 2023, 1:32am UTC](https://community.openai.com/t/confidence-score-for-prompt-response/132278/10 "2023-08-30T01:32:15Z")

</div>

> [@jwatte](#):
>
> Alright, I tried the actual API, and it does return the list of logprobs for each token in both the input and output.
> 
> Thus, you could use that information to calculate the “probability” of this particular sequence being chosen, and using that as some form of “confidence.

The API completion endpoint already includes a mechanism that does similar on the entire output: the “best\_of” parameter. Acting on the score instead of reporting it.

Like returning the logit probabilities themselves, it is not available for chat models such as gpt-4.

> **best\_of**
> 
> Generates `best_of` completions server-side and returns the “best” (the one with the highest log probability per token). Results cannot be streamed.
> 
> When used with `n`, `best_of` controls the number of candidate completions and `n` specifies how many to return – `best_of` must be greater than `n`.
> 
> **Note:** Because this parameter generates many completions, it can quickly consume your token quota. Use carefully and ensure that you have reasonable settings for `max_tokens` and `stop`.

This also doesn’t ensure entailment quality, just an overall “minimum deviation from the plan” generation that uses highest probability tokens. It is more for when you’ve allowed such creativity in token choices by not constraining temperature.

---

<div class="post-metadata">

**Author:** ![jwatte](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/jwatte/32/127705_2.png) [@jwatte](https://community.openai.com/u/jwatte)\
**Post date:** [September 1, 2023, 4:53pm UTC](https://community.openai.com/t/confidence-score-for-prompt-response/132278/11 "2023-09-01T16:53:43Z")

</div>

> [@\_j](#):
>
> the “best\_of” parameter

The “best\_of” parameter doesn’t actually return anything resembling an overall probability or confidence in the generated output, it controls the generation based on inferred probabilities, which is different.

Anyway, it sounds like we all agree:

1. The old completion API does return probabilities, which can be multiplied together after conversion (or added, and then exponentiated) to calculate “the probability of the generated text.”
2. This isn’t actually a very good measure of “confidence” anyway.
3. The new chat completion API doesn’t even give you the data.
4. LLMs don’t really have a concept of “the entire thing they generate” because that “entire thing” is a product of iteration that happens outside of the model itself; the model only considers each token at a time, so what the original request wants, doesn’t even exist.

---

<div class="post-metadata">

**Author:** ![EricGT](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/ericgt/32/20571_2.png) [@EricGT](https://community.openai.com/u/EricGT)\
**Post date:** [December 21, 2023, 10:20am UTC](https://community.openai.com/t/confidence-score-for-prompt-response/132278/12 "2023-12-21T10:20:45Z")

</div>


