Valid.operational AI assurance
Vendor-independent operational AI assurance

One evaluation harness across all operational AI systems.

Valid independently evaluates AI-generated military planning outputs across models, vendors, and versions.

Initial use case: per-COA assurance and decision-space assessment for AI-assisted planning.

operational planning assurance consoleRANK · ASSESS · DISTINGUISH · COVER
Valid dashboard ranking courses of action, separating individual COA assessment from decision-space assessment, and showing judge scores, assurance status, acceptability review, distinguishability, coverage, and duplicate detection.

Valid separates candidate ranking from assurance. Each COA is assessed for constraints, feasibility, suitability, acceptability, completeness, doctrine, and traceability, while the full option set is evaluated for distinguishability, coverage, and duplication.

Rank and assurance are separate: a top-ranked COA can still require human review, while the full option set can pass distinguishability and remain incomplete on coverage.

Multi-vendor operational AI

Multiple AI systems are generating planning content. Who independently assures the result?

Operational staffs increasingly receive planning outputs from different platforms, vendors, foundation models, software versions, and allied systems. Those outputs may conflict, rely on different sources, duplicate the same approach, or omit important alternatives. Without a common assurance layer, staffs are left to reconcile them manually before leaders decide.

Different methods and evidence

Outputs do not arrive on common ground.

Each system uses different methods, formats, assumptions, source practices, and documentation standards, making direct comparison difficult.

False decision-space diversity

More outputs do not guarantee better options.

Several models can produce superficially different plans that share the same assumptions, omissions, dependencies, or operational approach.

Independent assurance gap

Vendors cannot independently assure themselves.

Built-in checks may evaluate one system, but they do not provide a neutral comparison across the full set of planning outputs generated by different tools.

One independent assurance platform

Apply one assurance standard to each planning output and the full decision space.

Valid evaluates individual COAs, assesses the option set as a whole, and preserves the evidence and human adjudication behind every result.

Output assurance

Evaluate individual products.

Check doctrine, mission constraints, factual grounding, unsupported claims, evidence quality, and internal consistency.

Decision-space comparison

Compare systems and alternatives.

Apply consistent criteria across models, vendors, and versions while identifying material changes, redundancy, and missing options.

Evidence and adjudication

Preserve the review record.

Link findings to sources, doctrine, model metadata, staff decisions, revisions, and residual risk in one exportable assurance record.

COA assurance use case

More generated options do not always create a better decision space.

In one AI-assisted planning run, Valid evaluated 50 generated COAs, identified 37 distinct operational concepts, collapsed 13 duplicates or near-duplicates, and surfaced the smaller set that warranted further human review.

50
Generated COAs
37
Distinct concepts
13
Duplicates identified
9
Required human review
8
Advanced to the next planning step
2
Missing archetypes
What Valid surfaced

The set contained meaningful alternatives, but it also included redundant concepts and omitted two operational archetypes. Valid therefore evaluates both the quality of each COA and whether the commander is being presented with a sufficiently broad decision space.

MDMP assurance workflow

Validate the process before it becomes orders.

Valid supports an iterative staff workflow at defined decision points. Secure upload is the initial pilot method; connected assurance is the intended operating model. Valid does not replace MDMP, autonomously write the order, or approve a military decision.

1 · Receive planning outputs

Evaluate the current planning package.

During a pilot, staff securely upload mission-analysis products, COAs, wargaming results, comparison products, sources, metadata, or draft orders. Connected systems can deliver the same products directly as the workflow matures.

2 · Run assurance

Surface material planning weaknesses.

Check unsupported assumptions, doctrinal conflicts, weak traceability, redundant or missing COAs, inconsistent analysis, synchronization problems, and material changes across models or versions.

3 · Adjudicate and revise

Return findings to the responsible staff section.

Staff accept, reject, escalate, or assign each finding for revision. The responsible section updates the affected planning product and repeats analysis where required.

4 · Advance with evidence

Preserve the decision and assurance record.

Valid reruns affected checks and records what changed, what was resolved, and what residual risk was accepted before commander approval and order publication.

Recommended assurance gates

Gate 1: before COA approval, evaluate whether the decision space is distinct, complete, supportable, and consistently analyzed. Gate 2: before order publication, verify that the draft order faithfully translates the approved COA and that tasks, timing, resources, control measures, and annexes remain synchronized.

Valid findings dashboard showing an uploaded planning package, an evidence-backed blocker, recommended staff action, and human adjudication controls.
When Valid surfaces a material issue, staff can inspect the evidence, accept or reject the finding, assign it for revision, and preserve the adjudication record.
Transparent evaluation

Trace every ranking and finding to the evaluator evidence.

Valid preserves the evaluator trace behind each result, including repeated runs, COA-fit judgments, justification quality, doctrine coverage, constraint checks, unsupported details, and source conflicts. Reviewers can inspect why a result was produced instead of relying on an unexplained score.

Valid judge trace showing repeated evaluation runs, generator output, COA-fit judgment, justification-quality judgment, doctrine coverage, and an invented-detail finding.
The trace makes the evaluation inspectable rather than reducing the result to an unexplained composite score.
How Valid fits

Valid sits between generation and human approval.

Planning platforms and AI tools continue generating operational content. Valid provides a vendor-independent assurance layer that converts submitted or connected outputs into comparable assessments, evidence-backed findings, recommended staff actions, and an audit record.

Operational AI platforms, planning copilots, AI tools, allied systems, and human baselines feed outputs into Valid before human review and approval.
Start with secure uploads during a bounded pilot. After the workflow is proven, connect Valid directly to the systems producing planning outputs.
Start with a bounded pilot

Establish an independent assurance baseline.

Use historical, synthetic, or unclassified planning packages to test whether Valid surfaces material issues missed by the current review process, improves traceability, and adds acceptable review effort at a defined assurance gate.

What the team provides

A bounded planning package

  • One or more AI-assisted planning products
  • Relevant doctrine, constraints, and source references
  • A staff-authored or commander-approved baseline where available
What Valid returns

An adjudicated assurance record

  • Evidence-backed findings tied to affected products
  • Cross-model, version, revision, and decision-space comparisons
  • Recommended staff actions and an exportable audit record
How the pilot is judged

Clear operational success criteria

  • Material issues found beyond the existing review process
  • Finding precision and staff agreement after adjudication
  • Time to review, revise, rerun, and advance the product
Deployment path

Begin with secure planning-package uploads. The current product is designed to remain model- and vendor-agnostic. Connected workflows and secure APIs are roadmap capabilities. Secure deployment architecture is designed for future controlled-environment implementation; classified accreditation and additional operational workflows require customer-specific development and validation.

Recommended initial scope: three to five scenarios or planning packages, one defined assurance gate, named staff adjudicators, and pre-agreed measures for issue detection, review effort, and decision usefulness.

Pilot assessment

Get in touch for a pilot assessment.

Tell us what planning AI outputs you need to assess. We will follow up to discuss fit, scope, available data, and a bounded pilot design.

No sensitive or classified information.