NSGX

Evidence, in plain English

A better result needs a fair comparison.

The useful question is not whether a demo looks impressive. It is whether NSGX helps your system complete the same work with the quality, cost and speed you need.

Four questions behind the numbers

Source recall

Did it find what mattered?

Of the required source records, how many were retrieved? Finding all of them does not mean every returned record was useful.

Retrieval precision

How much was useful?

Of the records returned, how many were relevant? Read this alongside recall, not instead of it.

Answer or task quality

Was the work correct?

Did the final answer or completed task meet the agreed criteria? Correct retrieval and a correct answer are different outcomes.

End-to-end cost and time

What did it really cost?

Count model input, output and cached tokens, plus preparing, storing, updating and retrieving information. Measure the time to finish the task, not only the lookup.

Know what you are looking at

An illustration, a benchmark and a pilot are different.

Illustration
A scripted example explains a concept. It does not measure performance.
Internal benchmark
A repeatable test measures specified tasks and data. Its result applies to that test configuration.
Workload evaluation
An agreed comparison on representative work tests whether the result is useful in your environment.

This page explains the method; it does not present independently verified performance results. Ask for the dated run, dataset size, baseline, configuration and failures behind any benchmark claim. Results from one workload are not a guarantee for another.

Less text is useful. Lower total cost is the test.

Tokens are the units used to measure text a model processes. If a test uses one-third as many input tokens, that is about 67% fewer input tokens—not automatically 67% less total AI spend.

What a useful test report should include
  • The task, dataset size, version and date.
  • Your current baseline, including search, caching and compression already in use.
  • The same questions and success criteria for each comparison.
  • Retrieved-source quality, answer quality and failed attempts.
  • Input, output and cached tokens; retrieval time and task completion time.
  • Preparation, storage, update, retrieval and model costs, with the pricing assumptions.
  • Repeated runs and limitations—not only the best result.

Public reports can describe tasks, methods and aggregate results without publishing proprietary retrieval mechanics or customer records. Live tests may require a paid plan and provider usage charges; review their cost controls before running.