Not connected with the Agentix AI at agentixai.com. There is an established and entirely separate company trading under a very similar name at agentixai.com and at agentix.com. We have no relationship with it, no shared ownership and no shared personnel, and neither company endorses or is responsible for the other. If you were looking for them, you are on the wrong site.
Queensland, Australia
Did the agent do it, or did it only say so?
Agentix AI is building an evaluation harness that reads the end state of a machine rather than the closing message an agent wrote about itself. Nothing has shipped. This page sets out the seven stages the harness is being designed around, with the limit of each one written next to it.
The harness
Seven stages, each with its limit stated.
An agent that reports success is producing a sentence. Whether the sentence is true is a separate question, answerable only from the machine it was working on. These seven stages are how we intend to answer it, and each one is published with what it cannot tell you, because a measurement sold without its limits is a claim rather than a measurement.
-
01
Task specification
A task enters the harness as a recorded starting state plus a set of conditions that a program can evaluate against the finishing state. The prompt handed to the agent is an input to the run, not the definition of success. A condition that cannot be written as a check does not become a softer check, it keeps the task out of the set until somebody writes it properly.
What it measures
Whether the goal has been stated precisely enough that a third party could decide the outcome without reading the agent's own account of it.
What it cannot do
Say whether the task is worth doing, or whether it resembles the work you actually care about. A precise specification of the wrong thing is still the wrong thing, and no amount of measurement downstream repairs that.
-
02
Sealed start state
The run begins inside a disposable container built from a recorded image. Filesystem contents, environment variables, installed package versions, the clock offset and the outbound network policy are all fixed and hashed before the agent is permitted to act.
What it measures
Exactly what the agent started from, to the byte, so that a later disagreement about a result is a disagreement about evidence rather than about anybody's memory of how the box was set up.
What it cannot do
Seal a live third party service. A run that reaches one is marked unsealed in the record, and its replay figures are reported apart from the sealed ones rather than averaged in with them.
-
03
Transcript capture
Every tool call is written down in order with its arguments, its result, its exit status and its timing, alongside the model's own visible output. The transcript is append only and hashed when the run stops, so a later edit is detectable rather than merely discouraged.
What it measures
What the agent did, in the order it did it, including the calls that failed, the ones it retried, and the ones it made after it had already decided it was finished.
What it cannot do
Recover intent. A transcript shows moves and never reasons, and any tool that offers to read motivation out of a log is guessing with extra steps.
-
04
End state diff
When the run stops, the container is compared against the sealed image. Files added, changed and removed, rows touched in any local database, processes left running, ports left open and outbound requests attempted are all enumerated with hashes.
What it measures
What actually changed. This is the one part of the pipeline that does not depend on a single thing the agent said about itself, which is why the rest of the pipeline is arranged around it.
What it cannot do
Tell you whether a change was a good idea. It reports the delta and hands the judgement back to the conditions written in stage one.
-
05
Assertion run
The conditions written alongside the task are executed against the finished container. Each condition resolves to pass, fail or error, the output of every check is kept, and an error is never quietly counted as a pass because that is the failure mode which makes a harness worse than no harness.
What it measures
Whether the recorded end state satisfies the conditions the task named, with the evidence for each verdict stored next to the verdict.
What it cannot do
Catch a requirement nobody wrote down. Coverage of an assertion set is a property of the person who wrote the task, and the report says how many conditions were checked rather than implying the list was complete.
-
06
Claim reconciliation
The agent's closing message is split into discrete factual claims about what it did, and each claim is matched against the diff and the transcript. A claim resolves as supported, contradicted or unresolvable. This stage is the reason the other six exist.
What it measures
The distance between what an agent reported and what the machine shows. A run can satisfy every assertion and still close with a summary full of contradicted claims, and a harness that only prints a pass rate will never show you that.
What it cannot do
Resolve a claim too vague to be checked. Those are counted as unresolvable and published as their own figure, because folding them into the pass rate would flatter the agent, the harness and us at the same time.
-
07
Replay and record
The same task is run repeatedly from the same sealed state, and the outcomes are kept together in one record: inputs, image hash, transcript hashes, diffs, assertion results, the claim table and the software versions that produced all of it.
What it measures
Variance. How often the same agent, given the same task and the same starting state, arrives somewhere different, which is a number most published agent results quietly leave out.
What it cannot do
Make a stochastic system deterministic, and it is not a certification. A record is evidence somebody else can read and disagree with. It accredits nobody, including the company that produced it.
The gap
Pass rates are the wrong number on their own.
Two different questions
Whether an agent completed a task and whether an agent believes it completed a task are separate facts about the world, and they come apart often enough to be worth measuring on purpose. A benchmark that reports one number is answering the first question and quietly assuming the second.
The interesting failures are not the runs that fail loudly. They are the runs that finish with a confident summary describing work that the diff cannot find. Nothing in an ordinary pass or fail column records that a summary was wrong, so nobody counts it, so nobody improves it.
What the record would show
- Assertions passed, out of assertions written, with the count of conditions and not just the ratio.
- Claims supported by the diff, claims contradicted by the diff, claims that could not be resolved either way.
- Variance across repeated runs from one sealed starting state.
- Whether the run was sealed at all, since anything touching a live service is not reproducible.
Where this could be wrong
Splitting a closing message into checkable claims is itself a judgement, and a bad split produces a bad number. That step is the weakest part of the design, we know it, and the intended answer is to publish the split alongside the verdict so a reader can dispute it rather than take it on faith.
Scope
Four things this is not.
Written before there is a product, so that they cannot be quietly dropped after there is one.
No ranking of models
A harness that publishes a table of vendors becomes a marketing surface, and the incentive to keep the table interesting corrupts the measurement. The intended output is a record about one run, readable by the person who commissioned it.
No judgement about harm
Whether an action was dangerous, deceptive or improper is a question about values and context. The harness reports what changed and what was claimed. Deciding what that means is work for a person who understands the setting.
No badge to put on a page
We will not issue a mark, a seal or a score anyone can display as proof of anything. A record is evidence that can be read and argued with, which is the opposite of a badge.
Nothing has shipped
There is no product to buy, no early access list, no waiting list and no beta. The company was registered in 2026 and has released nothing. Everything above describes intent.
Status
The verifiable part.
Everything on this page that is checkable by a stranger, and nothing that is not. No clients, no users, no funding, no awards, no case studies, no shipped software.
- Legal name
- AGENTIX AI PTY LTD
- Entity type
- Australian proprietary company, limited by shares
- ACN
- 695 748 693
- ABN
- 24 695 748 693
- GST
- Not currently registered for GST
- State
- Queensland, Australia
- Registers
- The ACN is held by the Australian Securities and Investments Commission. The ABN and the GST position are published free of charge at abr.business.gov.au
- Service of documents
- The registered office recorded against ACN 695 748 693 at ASIC is the address with legal effect. We do not print a second address here that would not have that effect
- Trading record
- None. No product has been released, no service has been sold, and no revenue has been earned
AGENTIX AI PTY LTD does not hold ISO/IEC 27001 certification, a SOC 2 Type I or Type II report, an IRAP assessment or any other independent accreditation, and will not represent otherwise until one is genuinely held. There is no external audit of the harness, no penetration test, no insurance certificate to wave and no third party review of anything described on this site.
One address, no form.
Every route reaches the same inbox. Ordinary correspondence is answered within five business days, and a privacy request within thirty days under the Privacy Act 1988 (Cth).