Accuracy and coverage
Does the result match verifiable facts and cover the accessible elements of the task?
Evaluating AI within an application means observing how it supports a real task, what it misses and what the professional still needs to check.
The protocol covers an application version, an AI feature, users and operating conditions. Property inspection software, diagnostic assistance, hospitality tools or robot-embedded AI: each assignment starts with an actual supported task and the expertise it requires.
This pilot methodology still needs field validation. No actual results or scores from this protocol have been published. Access, experts, evidence and test arrangements are confirmed in the quote: remote Light, or Full with on-site observation.
Does the result match verifiable facts and cover the accessible elements of the task?
Are significant risks recognised, with an alert and a referral to the appropriate professional?
Does the application distinguish facts, assumptions and missing information, without inventing an answer?
Can each result be linked to the correct case, its source and the version tested?
Can an error be corrected during normal work, with the correction retained and verifiable?
Can the professional understand a suggestion, check or reject it, and take back control?
The pilot targets at least six relevant scenarios per application: at least three routine, two difficult and one critical. A frequent situation can also be critical, such as providing allergen information. An out-of-scope feature calls for a different suitable scenario; its absence is not a failure.
Eighteen candidate scenario sheets cover property, automotive and hospitality work. Examples include assigning photographs to the correct property, recognising diagnostic limitations and checking an ingredient. Dangerous situations are simulated outside live operations.
Scope and workload are agreed before commitment. Fewer than six relevant cases produce a reduced exploratory study without an index. Six cases do not provide a statistical guarantee.
Interface, hardware and data issues are distinguished from AI behaviour. Shared evidence is limited to authorised data, with agreed access and retention arrangements.
Eight sections connect observations to action: assignment, summary, test context, protocol, findings, actions, appendices and review. The report identifies completed cases, limitations, incidents and outstanding checks.
Its conclusion may recommend a supervised trial, limited use under documented conditions, corrections and retesting, advise against use within the tested scope, or state that no conclusion is possible. A serious blocker remains visible and cannot be offset by a favourable average.
The structured case set helps repeat tests and assess a corrected version. It complements professional judgement; comparing two applications requires compatible scopes and conditions.
An optional index is being prepared for a private pilot. It remains experimental, with no public score or badge. Contextualised results may be considered in V2, followed by monitoring in V3 once the methodology is validated and the client agrees.
No. It prepares a documented professional opinion within a specific scope. It does not certify the application or its overall compliance, and does not replace the operator’s decision.
When useful and feasible, a human-only reference can be compared with AI-assisted work under comparable conditions: information, equipment and task order. We examine quality, omissions and time spent checking and correcting. A small pilot does not establish a general productivity gain.
The NIST ARIA manual and AI RMF resources inform this work. Pilot dimensions, thresholds and rules are Huxpert design choices, without NIST certification or endorsement.
Describe the application, its users and the decision the evaluation should support.