Reading answers
What each field in a Drex response means, why answers are probabilities, and how to turn them into decisions.
Drex does not return a verdict. It returns a probability for each question, and your code turns that into a decision. This page explains each field in a response and the reasoning behind choosing thresholds, rounding, or using the whole distribution. For how to shape the questions, see Questions.
The response shape
A successful POST /v1/systemone returns HTTP 200 with one answer per question id. This is an example output for a support ticket. Numbers vary slightly per call.
{
"model": "drex-v1.0",
"answers": {
"category": {
"type": "choice",
"choice": "billing",
"confidence": 0.986,
"probabilities": { "billing": 0.9895, "technical": 0.005, "account": 0.0006, "other": 0.0049 }
},
"sentiment": {
"type": "score",
"score": 1.5306,
"legend": ["Calm", "Mildly annoyed", "Frustrated", "Angry"],
"probabilities": { "0": 0.1572, "1": 0.3516, "2": 0.2946, "3": 0.1966 },
"confidence": 0.7183
},
"wants_refund": {
"type": "noul",
"noul": 0.9868
}
},
"usage": { "input_tokens": 106, "output_tokens": 212 },
"evaluation_time_ms": 142.6,
"request_id": "req_<32 hex characters>"
}The state was "I was charged twice for my March invoice. Please refund one of the charges." Each answer repeats its type, so you can switch on it without looking up the question.
A noul answer is the probability of yes
noul is a number from 0 to 1. It is the model's probability that the answer to your question is yes. For the ticket above, "Is the customer asking for a refund?" came back at 0.9868. For "Does this describe a service outage?" on an SSO ticket that locked out 40 people, it came back at 0.7632. For "Is this review a bug report?" on a complaint that a page takes five seconds to open, it came back at 0.3503.
There is no built-in cutoff. You choose a threshold, and the right one depends on what each kind of mistake costs you.
- If acting on a false yes is cheap and missing a true yes is expensive, set a low threshold. Tagging a ticket
maybe_refundfor a human to glance at costs a few seconds. Missing a refund request costs a customer. - If acting on a false yes is expensive, set a high threshold. Auto-issuing a refund at 0.6 will refund people who did not ask for one.
- If both mistakes are expensive, use two thresholds. Act above the high one, ignore below the low one, and send the band in between to a person.
A threshold is a business decision, not a property of the model. Log the raw noul value with each decision, so you can move the threshold later and replay past traffic against the new one.
A choice answer is a label and a distribution
A choice answer has three fields.
choiceis the selected label, one of the keys you sent incriteria.probabilitieshas one entry per label. In our captures they sum to 1 within rounding. Thecategoryanswer above sums to 1.0000.confidenceis the average of the confidence (probability) scores the model produced for this answer. It measures how sure the model was overall, so it tracks the top probability but is a different number. Above,billinghas probability 0.9895 andconfidenceis 0.986. On another ticket,p0had probability 0.8966 andconfidence0.8621. Do not compute one from the other.
choice alone is enough when one label dominates. When the top two are close, probabilities tells you that the model could not separate them. A moderation question on a vague threat came back as flag at 0.48 with remove at 0.34. Code that reads only choice sees a clean flag. Code that reads probabilities can send that comment to a reviewer because a third of the mass sits on remove.
A score answer is an expected level
A score answer places the state on your ordered scale.
legendis the array of level labels you sent, so index 0 is the first label.probabilitiesis keyed by index as a string,"0"to"n-1", one entry per level.scoreis the expected level, the sum of each index times its probability. Forsentimentabove, 0 × 0.1572 + 1 × 0.3516 + 2 × 0.2946 + 3 × 0.1966 = 1.5306, thescorereturned. It can fall between two levels. In every capture on this page it did.confidenceis the average of the confidence (probability) scores the model produced for this answer. A lower value means the model was less sure.sentimentabove, with its mass spread over four levels, hasconfidence0.7183. The satisfaction rating below, with more than half its mass on one level, has 0.8863.
Rounding score to the nearest level loses information. 1.5306 rounds to 2, Frustrated. But the distribution puts 0.5088 of the mass on Calm and Mildly annoyed together, so "Frustrated" describes this ticket less than half the time. Three ways to read the same answer, each right for a different job:
| You want | Read | For sentiment above |
|---|---|---|
| One label for display | The index with the highest probability | Mildly annoyed (0.3516) |
| A number to average across many tickets | score | 1.5306 |
| The chance the customer is at least frustrated | probabilities["2"] + probabilities["3"] | 0.4912 |
The fractional score is the right input for an average or a trend line. Averaging rounded labels throws away the fraction every time. A second capture shows the other case. A satisfaction rating of "The new onboarding flow is fine, I guess." scored 2.173 on a five-level scale with 0.5769 of the mass on Neutral. There the argmax and the rounded score agree, and either is fine for display.
usage, evaluation_time_ms, and request_id
usage.input_tokens is the number of tokens the model read for state and questions. This is what Drex bills. usage.output_tokens is what the model produced and is not billed. Rates are on the pricing reference.
evaluation_time_ms is the time the model spent on your request. It is not your round trip. Network time, queueing, and Drex's own validation and rate limiting are on top of it. In the captures on this page the model took between 120 ms and 145 ms. Under load, Drex waits up to 55 seconds for the model before it gives up, so set your client timeout with that in mind. See Errors and retries.
request_id is a string like req_ followed by 32 hex characters. It is also sent as the x-request-id and x-typesafe-request-id headers, and it is present on error responses too. Log it with every call. When you write to support@nace.ai, include it, and Drex can join your report to every log line for that request.