Huxpert
Pilot methodology · v0.1.1

Real work scenarios.
Evidence for decisions.

Evaluating AI within an application means observing how it supports a real task, what it misses and what the professional still needs to check.

A specific question and a defined scope.

The protocol covers an application version, an AI feature, users and operating conditions. Property inspection software, diagnostic assistance, hospitality tools or robot-embedded AI: each assignment starts with an actual supported task and the expertise it requires.

This pilot methodology still needs field validation. No actual results or scores from this protocol have been published. Access, experts, evidence and test arrangements are confirmed in the quote: remote Light, or Full with on-site observation.

Six complementary dimensions.

Accuracy and coverage

Does the result match verifiable facts and cover the accessible elements of the task?

Safety and domain consequences

Are significant risks recognised, with an alert and a referral to the appropriate professional?

Uncertainty and limitations

Does the application distinguish facts, assumptions and missing information, without inventing an answer?

Traceability and evidence

Can each result be linked to the correct case, its source and the version tested?

Workflow and recoverability

Can an error be corrected during normal work, with the correction retained and verifiable?

Human control and clarity

Can the professional understand a suggestion, check or reject it, and take back control?

Routine, difficult and critical cases.

The pilot targets at least six relevant scenarios per application: at least three routine, two difficult and one critical. A frequent situation can also be critical, such as providing allergen information. An out-of-scope feature calls for a different suitable scenario; its absence is not a failure.

Eighteen candidate scenario sheets cover property, automotive and hospitality work. Examples include assigning photographs to the correct property, recognising diagnostic limitations and checking an ingredient. Dangerous situations are simulated outside live operations.

Scope and workload are agreed before commitment. Fewer than six relevant cases produce a reduced exploratory study without an index. Six cases do not provide a statistical guarantee.

How an evaluation is organised.

  1. Prepare. Record the version, conditions, criteria and professional reference before reviewing AI answers. Document evaluators’ expertise and relevant relationships.
  2. Observe. Retain authorised inputs, raw outputs and identified evidence. Separate application results from human corrections, including time spent checking and reworking.
  3. Review. The pilot plan calls for at least two executions of each critical case, followed by a second rating of critical evidence and a sample of other cases, without access to the first ratings. Disagreements are recorded and resolved.
  4. Report and retest. Explain discrepancies, their consequences and priority corrections. After an update, rerun affected cases under documented conditions.

Interface, hardware and data issues are distinguished from AI behaviour. Shared evidence is limited to authorised data, with agreed access and retention arrangements.

A report that supports a decision.

Eight sections connect observations to action: assignment, summary, test context, protocol, findings, actions, appendices and review. The report identifies completed cases, limitations, incidents and outstanding checks.

Its conclusion may recommend a supervised trial, limited use under documented conditions, corrections and retesting, advise against use within the tested scope, or state that no conclusion is possible. A serious blocker remains visible and cannot be offset by a favourable average.

Questions about the methodology.

Is this a benchmark?

The structured case set helps repeat tests and assess a corrected version. It complements professional judgement; comparing two applications requires compatible scopes and conditions.

Will there be a Hu-score?

An optional index is being prepared for a private pilot. It remains experimental, with no public score or badge. Contextualised results may be considered in V2, followed by monitoring in V3 once the methodology is validated and the client agrees.

Does this protocol provide certification?

No. It prepares a documented professional opinion within a specific scope. It does not certify the application or its overall compliance, and does not replace the operator’s decision.

What changes with human review?

When useful and feasible, a human-only reference can be compared with AI-assisted work under comparable conditions: information, equipment and task order. We examine quality, omissions and time spent checking and correcting. A small pilot does not establish a general productivity gain.

Methodological references

The NIST ARIA manual and AI RMF resources inform this work. Pilot dimensions, thresholds and rules are Huxpert design choices, without NIST certification or endorsement.

What task should your AI handle well?

Describe the application, its users and the decision the evaluation should support.