Source recall
Did it find what mattered?
Of the required source records, how many were retrieved? Finding all of them does not mean every returned record was useful.
Evidence, in plain English
The useful question is not whether a demo looks impressive. It is whether NSGX helps your system complete the same work with the quality, cost and speed you need.
Source recall
Of the required source records, how many were retrieved? Finding all of them does not mean every returned record was useful.
Retrieval precision
Of the records returned, how many were relevant? Read this alongside recall, not instead of it.
Answer or task quality
Did the final answer or completed task meet the agreed criteria? Correct retrieval and a correct answer are different outcomes.
End-to-end cost and time
Count model input, output and cached tokens, plus preparing, storing, updating and retrieving information. Measure the time to finish the task, not only the lookup.
Know what you are looking at
This page explains the method; it does not present independently verified performance results. Ask for the dated run, dataset size, baseline, configuration and failures behind any benchmark claim. Results from one workload are not a guarantee for another.
Tokens are the units used to measure text a model processes. If a test uses one-third as many input tokens, that is about 67% fewer input tokens—not automatically 67% less total AI spend.
Public reports can describe tasks, methods and aggregate results without publishing proprietary retrieval mechanics or customer records. Live tests may require a paid plan and provider usage charges; review their cost controls before running.