Confidently Wrong: What Four Frontier Models Got Wrong About Australian Tax Law
Your AI read the internet. Ours read what tax and legal professionals read. We benchmarked Lawpath Atlas against four frontier models and Perplexity on 30 recent Australian tax and legal changes, and published every score, including the one we think we got wrong.
Treasury’s Payday Super factsheet, last updated September 2024, says superannuation contributions must reach the fund within 7 calendar days of payday. The enacted law says seven business days. Alvarez & Marsal’s technical update flagged that shift as a critical change from the exposure draft, the kind of thing a payroll manager has to act on and a 2024 press release will never tell them. The government source is primary, well-ranked, and wrong. Professional commentary carried the corrected number because practitioners had to reconfigure payroll around it. That difference, between what a search finds and what a practitioner needs, is what this benchmark measures.
A small business owner does not know when an AI is guessing. They see a specific number, a specific date, a specific case citation, delivered with the same even confidence whether it is right or eighteen months out of date. Without a browser, the raw models in this benchmark were confidently wrong far more often than they were right. That is the failure mode we built Lawpath Atlas to close, and it is the one we set out to measure.
Atlas is the AI advisor behind Lawpath, answering legal, tax, compliance and accounting questions for over 650,000 Australian businesses. It works because a weekly knowledge ingestion pipeline keeps its Australian legal and tax knowledge current, rather than frozen at a training cutoff. We wanted a real number for how much that architecture actually buys a business owner, not just a marketing claim, so we built a ground-truth benchmark and are publishing the whole thing, including the parts that made us second-guess our own test.
The setup
Two rounds, six AI systems, 30 real Australian tax and legal questions, one primary judge model, and every score published.
We wrote the 30 questions from 84 Australian tax, superannuation, legal and accounting policy changes that Lawpath Atlas tracked from authoritative sources (William Buck, Andersen, PwC, TaxEd) over the eight weeks to 3 August 2026. That window was deliberate: it sits exactly where a model’s built-in knowledge is least reliable, so a wrong answer means the knowledge failed, not that the question was obscure. It is also a source of bias we address directly in the caveats below, because every question came from Atlas’s own tracked changes.
Questions were split into three tiers: low (one verifiable figure, date or citation), medium (a fact plus its mechanism), and high (interacting rules).
We ran two passes under the identical advisor system prompt and the identical structured rubric. Pass one, weights only: Claude Opus 5, GPT-5.6 Sol, GPT-5.6 Terra and Claude Sonnet 5 answered from model weights alone, with no retrieval and no tools, against Lawpath Atlas answering only from its continuously updated Australian knowledge base. Pass two, browser restored: the same four models answered again with their native web search switched back on (Anthropic’s server-side web_search for the Claude arms, OpenAI’s hosted web_search for the GPT arms), joined by Perplexity Sonar Pro. Lawpath Atlas’s pass-one answers carried forward unchanged, since it never searches. Perplexity’s pipeline was pinned to authoritative Australian domains, a constraint the other search arms did not have; if that hurt Perplexity’s recall, its scores below should be read as a floor, not a ceiling, for what an unconstrained Perplexity would score.
We graded 300 answers across both passes with zero request failures: 150 in pass one across five arms, and another 150 in pass two across five arms, since Atlas’s pass-one answers carried forward rather than being re-graded. Full marks require the figure, date and citation to match the known fact; a confidently stated wrong figure or date scores 0, never 1. Every answer was graded by an independent judge model (Kimi K3), and a separate robustness pass had every answer independently re-graded by a second frontier model from a different model family, under the identical rubric. Rankings and every headline gap held across both judges.
Pass one: what the models know without a browser
Two of the 30 questions, M7 and H6, are flagged as a likely error in our own ground truth, explained in full further down. We report the headline numbers on the clean 28 questions, and show the full 30-question data throughout for anyone who wants to check our working.

| Arm | Low | Med | High | Overall | Full marks (/28) | Partial | Declined | Confidently wrong | p50 latency |
|---|---|---|---|---|---|---|---|---|---|
| Lawpath Atlas | 2.00 | 2.00 | 1.78 | 1.93 | 26 | 2 | 0 | 0 | 0.8s* |
| Claude Opus 5 (raw) | 1.30 | 1.33 | 1.44 | 1.36 | 14 | 10 | 1 | 3 | 54s |
| GPT-5.6 Sol (raw) | 0.90 | 1.22 | 1.22 | 1.11 | 10 | 11 | 1 | 6 | 13s |
| GPT-5.6 Terra (raw) | 0.60 | 1.00 | 1.22 | 0.93 | 6 | 14 | 0 | 8 | 14s |
| Claude Sonnet 5 (raw) | 0.50 | 0.78 | 1.11 | 0.79 | 4 | 14 | 9 | 1 | 15s |
*Lawpath Atlas figure is retrieval-only p50. Raw arms had no retrieval and no tools to fall back on. “Declined” here counts only answers that were both explicitly declined and scored zero, a clean refusal rather than a guess; a handful of hedged answers elsewhere in the raw data were flagged as declines but still earned partial credit, which is a different thing.
With no tools at all, Lawpath Atlas fully answered 26 of 28 questions. The best raw frontier model, Claude Opus 5, managed 14. GPT-5.6 Sol answered 10 fully, GPT-5.6 Terra 6, and Claude Sonnet 5 only 4. Retrieval-grounded knowledge did not edge out a bigger raw model here. It more than doubled the next-best result.
Confident staleness, not ignorance
The more useful finding is not the accuracy gap itself but what sits underneath it. We graded two failure modes separately: declining to answer (the model explicitly flagged that it could not answer rather than guess, and scored zero for it) and confidently wrong (accuracy scored 0 without declining, meaning a stale or invented figure, date or citation was asserted as plain fact).

GPT-5.6 Terra was confidently wrong on 8 of 28 questions. GPT-5.6 Sol on 6. Lawpath Atlas, zero. The most cautious raw model, Claude Sonnet 5, largely avoided the trap by declining outright on 9 of 28 questions, and still only reached 4 full marks. Refusing to guess is safer than guessing wrong, but a small business owner who gets “I’m not sure” a third of the time will stop asking, which is its own kind of failure.
The dominant behaviour across the raw arms was not “I don’t know.” It was a specific, wrong number, stated with total conviction, and independently confirmed wrong under both judges. A small sample from this run:
| Question topic | What a raw model said | Correct current fact (verified) |
|---|---|---|
| Concessional super cap, FY2027 | Sol: “not yet determined, likely stays at $30,000” | $32,500 from 1 July 2026, up from $30,000 |
| Instant asset write-off from 1 Jul 2026 | Terra, Sol and Opus 5: “reverts to $1,000” | $20,000, made permanent from 1 July 2026 |
| Super guarantee charge percentage | Terra: “15%“ | Permanently 12% under Payday Super |
| Standard work-expense deduction | Terra: “$1,200” | $1,000, Sch. 4, Treasury Laws Amendment (Tax Reform No. 1) Bill 2026 |
| Bendel High Court citation | Terra, Sol: “[2025] HCA 15, 14 May 2025” | [2026] HCA 18, decided 10 June 2026 |
| Maximum super contribution base | Terra: “$260,000 p.a.” | $270,830 for FY2026-27, now an annual base |
Every one of these is the kind of answer a customer would act on directly: pay the wrong super contribution, claim the wrong deduction, cite the wrong case to the ATO. None of them came with a hedge.
Pass two: give every model a browser
Raw model weights are not how most people experience Claude or ChatGPT day to day. Real products bolt on web search, so we restored it for every model and added Perplexity Sonar Pro, a system built around search rather than retrofitted with it.

| Arm | Low | Med | High | Overall | Full (/28) | Partial | Confidently wrong | p50 | Searches/case |
|---|---|---|---|---|---|---|---|---|---|
| Lawpath Atlas | 2.00 | 2.00 | 1.78 | 1.93 | 26 | 2 | 0 | 0.8s | none |
| Claude Opus 5 + web search | 2.00 | 1.89 | 1.89 | 1.93 | 26 | 2 | 0 | 33s | 3.0 |
| Claude Sonnet 5 + web search | 2.00 | 1.89 | 1.78 | 1.89 | 25 | 3 | 0 | 24s | 2.3 |
| GPT-5.6 Terra + web search | 2.00 | 1.89 | 1.67 | 1.86 | 24 | 4 | 0 | 22s | 4.3 |
| GPT-5.6 Sol + web search | 2.00 | 1.89 | 1.67 | 1.86 | 24 | 4 | 0 | 27s | 3.6 |
| Perplexity Sonar Pro | 2.00 | 1.56 | 1.44 | 1.68 | 19 | 9 | 0 | 5.1s | native |
At n=28 per arm (ten low-tier questions, nine medium, nine high), the gap between 1.93 and 1.86 is one question. We do not read that as a ranking. The honest summary is that once every arm can search, Lawpath Atlas, Claude Opus 5, Claude Sonnet 5, GPT-5.6 Terra and GPT-5.6 Sol are statistically indistinguishable on accuracy, all landing within a question of each other, and Perplexity trails the field on medium and high-tier questions specifically. Latency, not accuracy, is what actually separates them, which is the point of the section below.
One pattern is worth calling out rather than treating as a strict win: Atlas scored a perfect 2.00 on the medium tier in both passes, the only arm to do so, against 1.89 for every search-enabled frontier model and 1.56 for Perplexity. Nine questions is not enough to call that statistically significant on its own. It is enough to notice that it is the same tier, every time, in both passes, and it lines up with a structural reason we lay out below.
Search also changes the shape of failure entirely. Every raw arm declined at least once; with search, declines vanish across the board, zero for every arm. Once the flagged M7 question is set aside, no arm was confidently wrong about anything, on any of the 28 remaining questions, once it could search.
Put the two side by side and the shift is obvious: the wall of grey “declined” and orange “confidently wrong” answers in pass one disappears entirely in pass two, replaced almost completely by full marks. The remainder that is not full marks lands as partial credit, a missing figure or instrument rather than an invented one. That is a real and welcome change in behaviour, and it is also the expensive option, which the next section puts a number on.
Speed, honestly
Atlas’s 0.8-second figure in the tables above is retrieval only. Every competing figure is end-to-end time to a graded answer, including the seconds a model spends composing the response after it has the facts, so the two numbers are not directly comparable as published.
Atlas’s own answer generation step, once retrieval hands it the facts, runs on a fine-tuned Gemma 4 model with a dedicated question-answering LoRA module, not a frontier-scale model, and takes roughly 10 seconds. Add that to the 0.8-second retrieval and Atlas’s realistic end-to-end answer time is an estimated 11 seconds: a sum of two separately measured medians, not a timed end-to-end run. Medians only add that way if the two components are correlated, so we are timing the full path properly before leaning on that number any harder.

That is still meaningfully faster: roughly 2 to 3 times faster than every search-enabled frontier model, without the extra variance that 2 to 4 web searches per question introduces. It is not faster than Perplexity, which answers in 5.1 seconds and beats Atlas on raw speed while trailing it by a wide margin on medium and high-tier accuracy. Fast and shallow is a real trade-off, not a solved problem, and we would rather say so than let the retrieval figure imply something the end-to-end number does not support.
We did not track cost per query in this benchmark run. Search multiplies both time and, almost certainly, API cost, since every search-enabled arm needed 2 to 4 web searches per question against Atlas’s zero. That is the other half of the “expensive option” argument, and it is currently unmeasured. A follow-up post will put a dollar figure next to the time figure.
Where the knowledge actually comes from
Everything above describes Atlas mechanically: curated weekly, already extracted, retrieved rather than reasoned. That undersells what the medium tier tests.
The Payday Super factsheet this post opened with is the pattern, not an exception, and that gap is what the medium tier measures. Atlas scored a perfect 2.00 on medium; search-enabled model arms sat at 1.89 and Perplexity at 1.56, on a 0 to 2 scale. Every arm, Atlas included, hit a clean 2.00 on the low tier. A generic search finds the headline figure reliably. It finds the mechanism, the instrument number and the assent status less reliably, because that detail lives in professional analysis rather than the announcement.
Three results from this run show the pattern.
The $270,830 maximum contribution base is meaningless without two further facts: it replaced a quarterly cap of $62,500, and it derives from the $32,500 concessional contributions cap divided by the 12% charge percentage. Search arms returned the figure and sometimes dropped the change.
Draft legislative instrument LI 2026/D3, the Out-of-Cycle Qualifying Earnings Determination, does not appear in most general Payday Super coverage. It determines which payments qualify for the extended contribution deadline rather than extending the deadline itself, and it shares its number with a separate live document, draft Law Companion Ruling LCR 2026/D3 on how the SG charge is calculated and assessed. An arm that returns the ruling instead of the instrument looks right and isn’t.
The $1,000 standard deduction and the $20,000 instant asset write-off both split the search arms, in opposite directions. The standard deduction is law: Schedule 4 of the Treasury Laws Amendment (Tax Reform No. 1) Act 2026, royal assent 26 June 2026. Arms that called it an announcement were reading April commentary. The permanent write-off is not law: announced in the 12 May 2026 Budget, sitting in the Treasury Laws Amendment (Tax Reform No. 2) Bill 2026, with the standing $1,000 threshold technically applying from 1 July 2026 until it passes. The $20,000 figure is legislated only to 30 June 2026, under the Strengthening Financial Systems and Other Measures Act 2025. Arms that called that measure announced were right; arms that saw $20,000 and called it law were reading last year’s extension. Professional updates track assent status explicitly, because practitioners cannot act on an announcement. That distinction, not a headline accuracy number, is the sharpest evidence in this benchmark for curated extraction over general search.
The question we think we got wrong
Methodology note
One question, M7, on the capital gains tax discount change, was scored 0 by every single arm in both passes: Lawpath Atlas, all four search-enabled models, and Perplexity Sonar Pro. All six systems, drawing on their own knowledge base or live web search independently, described the same companion measure (cost-base indexation plus a 30% minimum tax on real gains) rather than the specific fact the question asked for (a reduction in the CGT discount from 50% to 40%). When six independent systems, several of them searching the live web at the time of answering, converge on the same different answer, that is evidence the question's ground truth is wrong, not that six systems are wrong. Our working theory is that the case was written by conflating two separate measures from the same Budget announcement.
H6, a high-tier question, draws on the same disputed fact and is excluded alongside it. Every number in this post is computed on the remaining 28 questions. Including M7 and H6 at their actual scored values (30 questions, no exclusion) would show a 26 of 30 headline for Atlas and Claude Opus 5 with search, rather than 26 of 28; the ranking and every other conclusion is unchanged either way. We are flagging both for review rather than quietly excluding them from only the arms that benefit.
This matters for how you should read the rest of this benchmark. It is a genuine limitation of the test, not one we are hiding, and it is also a point in favour of grounded knowledge over raw guessing: every system that failed here failed for the same reason, a bad question, not five different kinds of wrong.
The gap is knowledge, not reasoning
Without search, it would be a more comfortable finding if the raw models were struggling with genuinely hard, multi-step legal reasoning. They were not. The pass-one gap was widest on the low tier, single-figure questions with one verifiable answer, where raw models scored between 0.50 and 1.30 against Atlas’s perfect 2.00. The gap narrowed on the high tier, where wording matters less than simply having the right number to reason from.
Search erases that low-tier gap entirely: every arm in pass two, Atlas included, scored a perfect 2.00 on every single low-tier question. With a browser, anyone can find a headline number. The contest moves to the medium tier, which is exactly where Atlas leads the whole field, for the structural reason described above.
A bigger model doesn’t close the gap. A browser does, at a cost.
Two separate levers were available to the raw models, and they behave differently. A better model buys real, measurable improvement: GPT-5.6 Terra (0.93) to GPT-5.6 Sol (1.11) to Claude Opus 5 (1.36). That still leaves more than half the questions not fully answered, at a median 54 seconds of thinking for the best of the three. A browser closes the accuracy gap almost entirely: the same Claude Opus 5, given web search, jumps to 1.93 and ties Lawpath Atlas outright, but it does not close the time gap. It moves it from 54 seconds to 33, still well behind Atlas’s roughly 11-second estimated end-to-end answer.
Both levers converge on the same conclusion: matching Atlas’s accuracy is achievable, with enough model or enough search time. Matching Atlas’s combination of accuracy and speed, at this scale, is not, because Atlas is not reasoning its way to the answer at request time at all. It is retrieving a fact that a weekly ingestion pipeline already found, extracted and verified before the question was ever asked.
What this means if you’re the one asking
If you ask Atlas about a recent change, such as the super cap or the instant asset write-off, the answer it gives should already carry the current figure, its effective date and the source it came from, because that is what the ingestion pipeline extracts before anyone asks. You should still check the citation if the decision is a large one, the same way you would ask a professional to point you to the source rather than take a verbal answer on faith. What this benchmark does not cover is evergreen, pre-cutoff Australian law, where the raw models score far higher and the gap this post describes mostly does not apply.
Caveats, read before quoting this
This compares systems, not pure model intelligence, across three configurations: retrieval-grounded generation (Atlas), raw model weights with no tools, and models paired with vendor web search.
The questions were drawn from the same 84 changes that Lawpath Atlas itself tracked, which means anything Atlas failed to ingest cannot appear in the question set. This benchmark measures Atlas’s retrieval and extraction quality conditional on ingestion succeeding, and structurally cannot penalise an ingestion miss. So read two Atlas numbers separately, because they measure different things even though the counts coincide. Ingestion recall: Atlas’s knowledge base held the complete correct fact for 26 of the 28 questions before pass one started, and that figure is the ceiling on everything else. Conversion: conditional on the fact being present, Atlas turned 26 of 26 into full marks, and its only two partial answers (H5 and H9) trace to the two questions where the base’s fact was incomplete. A cleaner version of this benchmark would draw its question set from an independent source chosen before checking what Atlas holds; that is an improvement we would make on a second run.
The questions are also deliberately adversarial to raw models: post-cutoff, recent-change questions, exactly where a retrieval or search system should win and a frozen model should struggle. On evergreen, pre-cutoff Australian law, we would expect the raw models to score far higher, and this benchmark says nothing about that regime.
Lawpath Atlas is not perfect either. Two answers scored partial credit at n=28, and the M7 miss sits in Atlas’s column the same as everyone else’s, at full weight, when the 30-question count is used instead.
Grading is by an LLM judge under a strict rubric. Figure, date and citation mismatches are objective and were consistent across both grading passes and both judges. Borderline phrasing judgements remain somewhat subjective, which is why we ran a second, independently-graded pass rather than reporting a single run.
We have not yet published a repository with the raw model outputs behind the “confidently wrong” table above. That is next, so the label does not have to be taken on faith.
Atlas extracts facts (caps, dates, instrument numbers, assent status) from professional firms’ technical updates rather than reproducing their commentary. The licensing position behind that is a fair question, and one we will answer directly in the comments rather than pre-argue here.
Why we published this
We have written before about the difference between a disclaimer and a safeguard. A disclaimer says “this might be wrong.” A safeguard is architecture that makes the system less likely to be confidently wrong in the first place, and honest about it when it might be.
This benchmark is that same posture applied to a number instead of a principle. We could have asserted that Lawpath Atlas has more current knowledge than a raw frontier model, stopped there, and looked good. Instead we ran a second pass that gave the raw models exactly the tool that would make our first result look less impressive, published the result where they caught up on accuracy, and flagged our own benchmark’s likely mistake in the open rather than quietly excluding the one question where everyone, including us, was wrong.
For the 650,000 Australian businesses that already trust Atlas, the point of this benchmark is not the accuracy tie. It is that the knowledge was already there, with its source and its currency attached, before anyone asked the question.
The complete data
Everything below is the full working behind the numbers above, including the two flagged questions in full, published for anyone who wants to check our grading rather than take our word for it.
Pass one, extended summary, all 30 questions, five arms
| Arm | Low | Med | High | Overall | Full (/30) | Partial | Miss | Declined | Confidently wrong |
|---|---|---|---|---|---|---|---|---|---|
| Lawpath Atlas | 2.00 | 1.80 | 1.70 | 1.83 | 26 | 3 | 1 | 0 | 1 |
| Claude Opus 5 (raw) | 1.30 | 1.20 | 1.30 | 1.27 | 14 | 10 | 6 | 7 | 3 |
| GPT-5.6 Sol (raw) | 0.90 | 1.10 | 1.10 | 1.03 | 10 | 11 | 9 | 3 | 6 |
| GPT-5.6 Terra (raw) | 0.60 | 0.90 | 1.10 | 0.87 | 6 | 14 | 10 | 2 | 8 |
| Claude Sonnet 5 (raw) | 0.50 | 0.70 | 1.00 | 0.73 | 4 | 14 | 12 | 15 | 1 |
“Declined” here counts every answer flagged as a decline regardless of score, including a handful that still earned partial credit by hedging rather than refusing outright, which is why declined and confidently wrong do not sum to miss on their own.
Pass one, every question, every arm
Each cell is the accuracy score out of 2. (dec) marks a question the model explicitly declined to answer rather than guess, including cases that still scored partial credit.
| Q | Tier | Lawpath Atlas | Claude Opus 5 | GPT-5.6 Sol | GPT-5.6 Terra | Claude Sonnet 5 |
|---|---|---|---|---|---|---|
| L1 | low | 2 | 1 (dec) | 0 | 2 | 0 (dec) |
| L2 | low | 2 | 2 | 2 | 2 | 0 |
| L3 | low | 2 | 1 (dec) | 0 (dec) | 0 | 0 (dec) |
| L4 | low | 2 | 2 | 2 | 1 | 1 (dec) |
| L5 | low | 2 | 1 (dec) | 0 | 0 | 1 (dec) |
| L6 | low | 2 | 2 | 2 | 0 | 2 |
| L7 | low | 2 | 0 | 0 | 0 | 0 (dec) |
| L8 | low | 2 | 2 | 1 | 1 | 1 |
| L9 | low | 2 | 0 | 0 | 0 | 0 (dec) |
| L10 | low | 2 | 2 | 2 | 0 | 0 (dec) |
| M1 | medium | 2 | 2 | 2 | 1 | 1 |
| M2 | medium | 2 | 1 (dec) | 1 | 1 | 1 (dec) |
| M3 | medium | 2 | 2 | 2 | 1 | 1 |
| M4 | medium | 2 | 1 | 1 | 1 | 0 (dec) |
| M5 | medium | 2 | 2 | 1 | 1 | 0 (dec) |
| M6 | medium | 2 | 1 | 1 | 1 | 1 |
| M7 | medium | 0 | 0 (dec) | 0 (dec) | 0 (dec) | 0 (dec) |
| M8 | medium | 2 | 0 (dec) | 0 | 0 | 0 (dec) |
| M9 | medium | 2 | 1 | 1 | 1 | 1 |
| M10 | medium | 2 | 2 | 2 | 2 | 2 |
| H1 | high | 2 | 2 | 2 | 2 | 1 |
| H2 | high | 2 | 2 | 1 | 2 | 2 |
| H3 | high | 2 | 0 | 0 | 0 | 1 (dec) |
| H4 | high | 2 | 2 | 2 | 2 | 1 |
| H5 | high | 1 | 1 | 1 | 1 | 1 |
| H6 | high | 1 | 0 (dec) | 0 (dec) | 0 (dec) | 0 (dec) |
| H7 | high | 2 | 1 | 1 | 1 | 0 (dec) |
| H8 | high | 2 | 2 | 2 | 1 | 2 |
| H9 | high | 1 | 2 | 1 | 1 | 1 |
| H10 | high | 2 | 1 | 1 | 1 | 1 |
M7 and H6, both flagged, are excluded from every headline number above but shown here at full weight.
Pass two, every question, every arm, search restored
| Q | Tier | Lawpath Atlas | Opus 5 + search | Sonnet 5 + search | Terra + search | Sol + search | Perplexity |
|---|---|---|---|---|---|---|---|
| L1 | low | 2 | 2 | 2 | 2 | 2 | 2 |
| L2 | low | 2 | 2 | 2 | 2 | 2 | 2 |
| L3 | low | 2 | 2 | 2 | 2 | 2 | 2 |
| L4 | low | 2 | 2 | 2 | 2 | 2 | 2 |
| L5 | low | 2 | 2 | 2 | 2 | 2 | 2 |
| L6 | low | 2 | 2 | 2 | 2 | 2 | 2 |
| L7 | low | 2 | 2 | 2 | 2 | 2 | 2 |
| L8 | low | 2 | 2 | 2 | 2 | 2 | 2 |
| L9 | low | 2 | 2 | 2 | 2 | 2 | 2 |
| L10 | low | 2 | 2 | 2 | 2 | 2 | 2 |
| M1 | medium | 2 | 2 | 2 | 2 | 2 | 2 |
| M2 | medium | 2 | 2 | 2 | 2 | 2 | 1 |
| M3 | medium | 2 | 2 | 2 | 1 | 2 | 1 |
| M4 | medium | 2 | 1 | 1 | 2 | 2 | 2 |
| M5 | medium | 2 | 2 | 2 | 2 | 2 | 1 |
| M6 | medium | 2 | 2 | 2 | 2 | 2 | 2 |
| M7 | medium | 0 | 0 | 0 | 0 | 0 | 0 |
| M8 | medium | 2 | 2 | 2 | 2 | 2 | 2 |
| M9 | medium | 2 | 2 | 2 | 2 | 1 | 1 |
| M10 | medium | 2 | 2 | 2 | 2 | 2 | 2 |
| H1 | high | 2 | 2 | 2 | 2 | 2 | 2 |
| H2 | high | 2 | 2 | 2 | 1 | 1 | 2 |
| H3 | high | 2 | 2 | 2 | 2 | 2 | 2 |
| H4 | high | 2 | 2 | 2 | 2 | 2 | 1 |
| H5 | high | 1 | 2 | 2 | 2 | 2 | 1 |
| H6 | high | 1 | 1 | 1 | 1 | 1 | 1 |
| H7 | high | 2 | 1 | 1 | 1 | 1 | 1 |
| H8 | high | 2 | 2 | 2 | 2 | 2 | 2 |
| H9 | high | 1 | 2 | 1 | 1 | 1 | 1 |
| H10 | high | 2 | 2 | 2 | 2 | 2 | 1 |
Notice how uniform the low tier becomes once everyone can search: a clean sweep of 2s. The differentiation that remains is almost entirely in the medium and high tiers, concentrated in H6 (tied to the flagged M7 issue) and H3 (where Claude Opus 5 with search uniquely reaches full marks by correctly identifying that corporate beneficiaries receive no franking credit).
Every confidently wrong answer, pass one
Every pass-one answer, across all five arms, that was stated as plain fact, with no hedge and no decline, and scored 0. In pass two, this list collapses to a single shared entry, M7, for every arm.
| Q | Arm | What it said | Correct fact |
|---|---|---|---|
| L1 | Sol | ”FY2027 cap not yet determined, likely stays at $30,000” | $32,500 from 1 July 2026 |
| L2 | Sonnet 5 | ”7 calendar days (not 7 business days)“ | 7 business days from each payday |
| L3 | Terra | ”$260,000 p.a. ($65,000 x 4)“ | $270,830 annual for FY2026-27 |
| L5 | Terra | ”[2025] HCA 15, 3 June 2025” | [2026] HCA 18, 10 June 2026 |
| L5 | Sol | ”[2025] HCA 15, 14 May 2025” | [2026] HCA 18, 10 June 2026 |
| L6 | Terra | ”$1,200, a Coalition proposal” | $1,000, Tax Reform No. 1 Bill 2026 |
| L7 | Terra | ”reverts to $1,000” | $20,000, permanent from 1 July 2026 |
| L7 | Sol | ”scheduled to revert to $1,000” | $20,000, permanent from 1 July 2026 |
| L7 | Opus 5 | ”the threshold reverts to $1,000” | $20,000, permanent from 1 July 2026 |
| L9 | Terra | ”30% from 1 July 2019, former Coalition policy, not enacted” | current Government proposal, 30% from 1 July 2028 |
| L9 | Sol | ”2017 ALP opposition policy, 1 July 2019” | current Government proposal, 30% from 1 July 2028 |
| L9 | Opus 5 | ”there is no current Australian Government proposal” | a proposal exists, 30% from 1 July 2028 |
| L10 | Terra | ”SGC percentage is set at 15%“ | permanently 12% |
| M7 | Lawpath Atlas | described the companion measure (indexation and a 30% minimum tax on real gains) but omitted the asked-for discount change | flagged as a likely case-quality issue, see above |
| M8 | Terra | ”five-year transition period” | three years from 1 July 2027 |
| M8 | Sol | ”two-year transitional period” | three years from 1 July 2027 |
| H3 | Terra | analysed a top-up-with-franking model | corporates get no credit, producing double taxation |
| H3 | Sol | analysed the 2017 opposition policy | corporates get no credit, producing double taxation |
| H3 | Opus 5 | ”no such measure in Australian law or before Parliament” | corporates get no credit, producing double taxation |
The 30 questions and their ground truth
| Q | Tier | Question | Known correct fact |
|---|---|---|---|
| L1 | low | What is the concessional superannuation contributions cap for FY2027 in Australia? | $32,500 from 1 July 2026, up from $30,000 in FY2026. |
| L2 | low | Under Payday Super in Australia, how many days does an employer have to get super contributions to the fund after each payday? | Within 7 business days of each payday, from 1 July 2026. |
| L3 | low | What is the annual maximum superannuation contribution base for FY2026-27 in Australia? | $270,830 for FY2027, replacing the quarterly base ($62,500 per quarter). |
| L4 | low | When does the ATO Small Business Superannuation Clearing House shut down permanently? | 30 June 2026. It closed to new users on 1 October 2025. |
| L5 | low | What is the citation and decision date for the High Court decision on unpaid present entitlements and Division 7A involving Bendel? | Commissioner of Taxation v Bendel [2026] HCA 18, decided 10 June 2026. |
| L6 | low | What is the amount of the proposed standard deduction for work-related expenses in Australia? | $1,000, under Schedule 4 of the Treasury Laws Amendment (Tax Reform No. 1) Bill 2026. |
| L7 | low | What is the instant asset write-off threshold for Australian small business from 1 July 2026? | $20,000, made permanent from 1 July 2026. |
| L8 | low | What superannuation balance threshold triggers the Division 296 tax in Australia and when does it start? | Balances above $3 million, from 1 July 2026. |
| L9 | low | What rate of minimum tax on discretionary trust income has the Australian Government proposed, and from when? | 30% minimum tax at the trustee level, proposed from 1 July 2028. |
| L10 | low | Under the Payday Super reforms, what is the superannuation guarantee charge percentage set at? | Permanently set at 12%. |
| M1 | medium | What replaces Ordinary Time Earnings as the superannuation guarantee calculation base under Australia’s Payday Super reforms, and what does it include? | Qualifying Earnings (QE) from 1 July 2026. Includes OTE, all commissions, directors fees, salary sacrificed amounts, and payments to contractors paid mainly for labour. Excludes overtime, expense reimbursements, certain allowances and specific termination payments. |
| M2 | medium | What did the High Court decide in the Bendel case about unpaid present entitlements owed by a trust to a corporate beneficiary? | A UPE does not automatically constitute a loan under s 109D(3) of Division 7A. Inaction by the corporate beneficiary is not financial accommodation. Decided 5 to 2, overturning the Commissioner position held since the 2010 financial year. |
| M3 | medium | How does the superannuation guarantee charge change under Australia’s Payday Super reforms? | SGC becomes tax-deductible and is assessed by the ATO rather than self-assessed. Includes shortfall, daily compounding interest at GIC rates, and an administrative uplift of up to 60%. Non-deductible penalties of 25% or 50% may apply per payday. |
| M4 | medium | What is the ATO transitional compliance approach for the first year of Payday Super in Australia? | PCG 2026/1, a risk-based approach for 1 July 2026 to 30 June 2027. Genuine efforts and prompt correction are low risk; unresolved shortfalls after 28 days from quarter end are high risk. A 12-month STP transition for QE data applies until 1 July 2027. |
| M5 | medium | What relief applies to out-of-cycle payments like bonuses and commissions under Australia’s Payday Super rules? | Draft legislative instrument LI 2026/D3 extends the deadline to 7 business days after the next regular payday rather than 7 days from the off-cycle payment. Termination payments are excluded. |
| M6 | medium | How does the proposed $1,000 standard work-related expense deduction interact with actual claimed expenses in Australia? | Where actual claims are less than $1,000, the standard deduction is reduced by the amount claimed so the total does not exceed $1,000. Where actual claims exceed $1,000, the standard deduction is reduced to nil. |
| M7 | medium | What change to the capital gains tax discount was proposed in the 2026-27 Australian Federal Budget? | Reduction from 50% to 40% for individuals above a specified taxable income threshold, increasing the taxable portion of gains for affected taxpayers. Flagged: every arm in both passes independently described a different Budget measure instead; treat this ground truth with caution pending review. |
| M8 | medium | What rollover relief is proposed for assets moved out of Australian discretionary trusts, and over what period? | Three-year rollover relief from 1 July 2027, allowing transfer of assets from discretionary trusts into companies or fixed trusts without triggering CGT or income tax. |
| M9 | medium | How does the superannuation guarantee opt-out work for Australian high income earners with multiple employers? | High income earners with multiple employers can apply for an SG exemption certificate for nominated employers where income would exceed the concessional cap ($32,500 for FY2027). The annual maximum contribution base means certificates can be granted even without concurrent employment in a quarter. |
| M10 | medium | What are the Australian public country-by-country reporting obligations for multinationals and when is the first lodgement due? | Public CBC reporting rules apply from 1 July 2024; first lodgement deadline for 30 June year-end entities is 30 June 2026. |
| H1 | high | After the High Court Bendel decision, what anti-avoidance provisions still apply to unpaid present entitlements from Australian trusts to corporate beneficiaries? | Subdivision EA, section 100A and Part IVA still apply. Existing arrangements should be reviewed before making structural changes; the decision does not make UPE arrangements unconditionally safe. |
| H2 | high | Under Australia’s Payday Super regime, what is the maximum number of penalty events a weekly-payroll employer could face in a year and why? | Up to 52 penalty events per year, because SGC is assessed per payday rather than quarterly, so each weekly pay cycle is a separate potential contravention. |
| H3 | high | How does the proposed 30% minimum tax on Australian discretionary trust income affect distributions to corporate beneficiaries specifically? | Non-corporate beneficiaries receive a non-refundable credit, but corporate beneficiaries receive no credit, effectively producing double taxation and diminishing the role of corporate beneficiaries (bucket companies) in trust distribution planning. |
| H4 | high | What should an Australian private group with a discretionary trust and corporate beneficiary review before 30 June 2026 following the Bendel decision? | Trust deeds, existing Division 7A loan agreements and accounting treatments. Groups may no longer need to place UPEs on complying Division 7A loan terms, potentially retaining funds in trusts at the company tax rate, but Subdivision EA, s 100A and Part IVA still apply. |
| H5 | high | What are all the changes an Australian employer must make to payroll for Payday Super from 1 July 2026? | Pay SG each pay cycle with receipt by the fund within 7 business days; calculate on Qualifying Earnings rather than OTE; apply an annual maximum contribution base of $270,830 instead of quarterly; SG charge percentage fixed at 12%; SGC assessed per payday by the ATO with up to 60% administrative uplift. |
| H6 | high | How do the proposed Australian CGT discount reduction and the shift to cost base indexation interact for individual investors? | The Budget proposes reducing the CGT discount from 50% to 40% for higher-income individuals, alongside a shift from the 50% discount to a cost base indexation model. Investors should weigh realising gains before commencement against market conditions. Flagged alongside M7, see above. |
| H7 | high | What recent Australian decisions have found workers to be employees for superannuation guarantee purposes despite contractor arrangements? | The Balmain Dental Clinic case (Administrative Review Tribunal) ruled a dental practitioner an employee for SG purposes; a Federal Court decision found a long-serving consultant to be a common law employee. Draft SGR 2026/D1 addresses employer identification in tripartite working arrangements. |
| H8 | high | What are the key GST margin scheme rules Australian property developers need to consider on acquisition eligibility and cost base? | Call option fees are excluded from the cost base; going concern and farmland acquisitions attract a look-through rule; amalgamated development sites have specific treatment; acquisition eligibility governs whether the scheme can be applied at all. |
| H9 | high | What changes have been made to the Australian R&D Tax Incentive regarding excluded activities and transparency reporting? | Tobacco and gambling activities are excluded from the R&D Tax Incentive from 1 July 2025. The R&D Tax Incentive Transparency Report for 2023-24 is to be published in September 2026. |
| H10 | high | What are the Australian director identification number obligations and what are the consequences of not having one? | Director ID is a strict liability criminal offence provision under s 1272C(1) of the Corporations Act 2001, in force from 5 April 2022, with phased application deadlines (30 November 2022 for pre-existing directors). Prosecutions have followed; the Business Registries Stabilisation Bill continues reform in this area. |
Lawpath Atlas is available for all Lawpath customers. To learn more, visit lawpath.com.au/lawpath-atlas.