A full cognitive warfare cycle, run under alliance conditions.
NATO Allied Command Transformation asked for an agentic AI that could run the whole counter-campaign cycle against a hostile information operation. We built it against a live one, and it was selected as one of ten solutions worldwide.
The mission
NATO Allied Command Transformation opened Innovation Challenge 2026-1, Agentic AI for Cognitive Warfare, under RFIP-ACT-SACT-26-38. The request named the work cycle it wanted closed: understanding, planning, coordination, delivery, assessment. It named who the campaigns would address, an adversary's defense and security apparatus and its military commanders. And it named the six effects a submission had to be able to produce against them.
Accepting a mission in that shape commits you to the whole cycle. A tool that only reads the environment answers the first stage and abandons the other four. A tool that only drafts answers the fourth and has no basis for any of it.
Weaken military strength and disrupt support networks
Damage the credibility and authority of senior leadership
Slow decisions and complicate the rollout of communications
Fragment internal unity and erode confidence
Impair effectiveness in reaching and influencing target audiences
Create false perceptions of the alliance's plans and operational intent
The problem
Counter-campaign work in this domain is bounded by a single thing: nobody can say in advance what a message will do to the people who receive it. Without that, planning is opinion, delivery is volume, and assessment measures activity instead of effect. Every stage of the cycle inherits the same blind spot.
The answer
Freya Sense models the effect itself, in both directions. Given a message, it projects how a defined audience will respond. Given an observed response, it infers what the message was built to do. That single capability is what lets one system carry all five stages, because every stage is asking a version of the same question.
The reasoning is inspectable, and it does not live inside a language model. The behavioral judgment is deterministic. Generative AI phrases what has already been decided, and never decides anything.
What had to be true
The full cycle runs end to end, with no stage handed to a separate tool or a separate vendor
Every campaign is calibrated to a declared effect on a declared audience, before it is drafted
Every decision is attributable to a person, by role and by timestamp
Assessment measures what happened against what was predicted, and the difference feeds the next cycle
The system operates where the data cannot leave, including inside a national enclave with no external connection
What we delivered
A concept paper answering the challenge against all six named effects and the full five-stage cycle
A technical summary naming every component, its maturity, and the three deployment topologies
A costed delivery plan across three phases, with the classified pathway and sovereign hosting identified
A recorded demonstration running the whole cycle against a live adversary operation, with an unforeseen variable injected
The conditions
NATO's Principles of Responsible Use for AI, PO(2024)0199, are the operating envelope in this domain rather than a compliance annex. They were answered one by one.
Rules of engagement and ethical thresholds are configuration, set by an administrator, rather than assumptions baked into the tool.
Approvals are role based, and every decision lands in an immutable audit log against a person and a time.
The causal model is whitebox, and every proposal carries the reasoning chain that produced it, back to the evidence.
Deterministic first, with failure made visible rather than smoothed over.
Kill switches at campaign level and at system level, and analyst override that is absolute.
Source-class weighting, ensemble classification, and a closed vocabulary constraining anything a language model is allowed to emit.
Three conditions the work imposed on itself, which matter more than the checklist.
Every variant stages for approval, and transmission is a separate action taken by an authorized person on an approved message.
The two roles are structurally separate, and no configuration collapses them.
What was observed was real, drawn from around 260 open-source items across three independent inventories. What was done in response was simulated: real tool behavior on real inputs, producing real proposals and real drafted text, with nothing transmitted.
Interoperability was treated the same way, against STANAG 4774 and 4778 for labelling and binding, and AQAP 2110 and 2210 for the assurance regime.
The result
Freya Sense was selected as one of ten companies worldwide, from an international field, and carried to Pitch Day at Rennes in May 2026.
What that selection says is narrow and worth stating precisely. An alliance body that had defined the problem judged this approach relevant, useful and feasible against its own five-stage cycle and its own six effects. That is the state of the art as it stood at submission. The system has moved since.
How it ran
The demonstration ran against a narrative operation live in the environment between 23 March and 17 April 2026.
Russian-origin drones penetrated NATO airspace three times inside that window: once in Lithuania on 23 March, and twice in Romania, on 25 and 26 March and again on 17 April. Kremlin-aligned media used those incidents as raw material for a single claim, that Lithuania, Latvia and Estonia had opened their airspace to Ukrainian drones attacking Russia. The claim originated on a Russian Telegram channel, spread through six Kremlin-aligned outlets within 48 hours, took state endorsement at a foreign ministry briefing on 6 April, and escalated on 16 April to an invocation of Article 51 of the UN Charter against the Baltic states.
The operation had a structural weakness, and it was geographic. The Baltic framing did not fit the incidents, because two of the three happened in Romania. Even pro-Kremlin commentators were skeptical of the claim, and the official escalation did not close that fracture.
Reading the environment produced the surge and the coordination signature. Reading it in reverse produced the intent: what behavior the adversary messaging was built to produce, in whom, and how hard. Four counter-campaigns were proposed in parallel.
Amplify the allied counter-messaging that was not reaching Russian-language audiences
Widen the fracture the adversary's own commentators had opened
Reframe against the physical record of where the drones actually came down
Apply targeted pressure at the point driving the escalation
Assessment then compared what the system predicted against what the adversary narrative actually did next. That comparison is the debrief, and it is what makes the next cycle sharper than the last.
The alliance set the mission, the conditions and the standard. What this account shows is that the cycle closes under them, against a real operation.
Countering a live operation against a national defense program.
A European presidential administration needed to answer an information operation that had already taken hold. What it asked for was messages that would land on the people the operation had already reached.
The mission
A national counter-disinformation program, running under the high patronage of a European presidential administration, brought two narratives it was facing. Both attacked the financing instrument behind a strategic defense program. Both had already gained traction inside a real audience in the capital.
What was asked for was a diagnosis of the effect each narrative was built to produce, and counter-narratives calibrated to the audience that had already absorbed them, in a form the program's own creative teams could take straight into production.
Accepting that commits you to something specific: whatever you hand over has to work in someone else's hands, on a deadline set by an operation that is still running.
The problem
Institutions that answer to a public can see a narrative spreading. What they cannot see is what it is doing to the people it reaches, or which of their own answers will move those people rather than the people who already agree. So the standing response is volume and speed, against an adversary who is constrained by neither.
The answer
Diagnose the effect, then calibrate to it. Each narrative was decomposed to the emotional and cognitive levers it was pulling, and to the points where it was structurally weak. Counter-narratives were then written against an existing behavioral model of that specific capital-city audience, so they arrived in the language and the register that audience responds to.
This changes what a communications team is deciding. They stop choosing between messages on instinct and start choosing against a stated effect.
What had to be true
Each narrative understood by the behavior it was engineered to produce rather than by the claims on its surface
The weak points of each narrative identified and written down
Counter-narratives calibrated to a real audience model rather than to a general public
Everything handed over in a form the client's own creative and advertising flow could use without us
What we delivered
A written diagnosis for each narrative: the effect it sought, the factors it used, and its points of vulnerability
Three counter-narratives per narrative, six in total, calibrated to the audience model and ready for production
A strategy workshop at the start, a debriefing workshop on the diagnosis, and a closing workshop on what the counter-narratives produced
The conditions
On one side of this work the requirement is that the system can attack. On this side the requirement is that it never does.
The integrity of the audience's own thinking is the boundary. The work brings what is true to the surface and takes the heat out of what is not. It does not move people by exploiting them, and no deliverable was written to.
Publishing stays with the client. Everything handed over is scaffolding for their creative flow. We never transmitted, never posted, and never held an account.
Public data only.
The audience model came first and was already built. Nothing was inferred about individuals in order to do this work.
Every counter-narrative was a proposal to a person who could reject it, and some were.
The result
The counter-narratives were adopted into the program's own production flow.
The work enabled a national team to produce fourteen counter-messages in less than a day, which materially accelerated its response. The videos built from that support reached roughly twice as many people as the operation they were countering.
How it ran
Two narratives in. A workshop to agree which two, what the program was trying to protect, and how success would be measured. Then the diagnosis, then a debriefing workshop where the diagnosis was argued with the people who owned the problem, then drafting, then a closing session on what the counter-narratives had produced and what to do next.
The debrief mattered more here than the diagnosis. A counter-narrative that a client's own creative team will not use is worth nothing, and the only way to know is to sit in the room while they take it apart.
The loop and the instruments are the ones the defense engagement uses. What changes is the rules of engagement. Beyond that boundary the work is a sword; here it is a glove.