Banking Technology

Access control, authentication and unsafe failure: what our AI security reviews found

Codex Daybreak Blue found reported vulnerabilities involving access control, authentication, sensitive information and unsafe failure handling in our AI-built API product.

RTD

RTD Team

Run-True Decision

Abstract editorial illustration of security review

Codex Daybreak Blue found reported vulnerabilities involving access control, authentication, sensitive information and unsafe failure handling in our AI-built API product. Among findings classified as security weaknesses, 26% were rated High and 51% Medium. Broken Access Control was the largest OWASP Top 10:2025 category, at 27%. These are useful priorities for a security review. They are not estimates of how often an attacker could compromise the system.

The study covers more than two hundred distinct reported findings.

Not all reported findings established a security weakness. The mapping classified 79% that way and excluded 21% from the security tables. The excluded group includes test quality, documentation and correctness concerns. Those concerns still matter. Counting them all as demonstrated vulnerabilities would make the headline stronger than the evidence.

Our decision software had no production customers during this period. The issues were found before any customer used the system. As of October 4, 2026, 82% of the security-classified findings had a merged scoped fix. That is not a claim that every issue was fixed, that each fix was deployed, or that the system had become safe in every operating condition.

Severity: reported ratings, not CVSS

The security-only distribution was Critical 0%, High 26%, Medium 51%, Low 20% and unrated 3%. No Critical finding was reported by Codex Daybreak Blue in this corpus. The labels are the reviewer's own ratings. They are not CVSS scores. CVSS, the Common Vulnerability Scoring System, is a standard way to describe vulnerability severity. We did not assign those standardized scores.

Reported severity among security-classified findings
Reported severity among security-classified findings

A High label helps prioritize investigation, but it is not a measured likelihood of exploitation. An unrated finding is not a Low finding. The recorded distribution reflects selected review targets and the evidence available to the reviewer. It should not be read as an inventory of every weakness that existed in the product.

OWASP categories: the API view and the wider application view

The OWASP API Security Top 10:2023 view puts Broken Authentication at 17% of security findings. Broken Object Level Authorization accounts for 7%, and Broken Object Property Level Authorization for 7%. The matching rounded authorization shares are not equal before rounding. These categories ask related but different questions: whether identity is reliable, whether that identity may access an object, and whether it may read or change the object's properties.

The API chart uses all security findings. A further 57% has no direct category in that API list. This does not mean those findings were safe or unrelated to security. The study also includes local tooling, operational configuration, logging and service-availability weaknesses. Forcing all of those into an API category would suggest a precision the classification does not support.

OWASP API Security Top 10:2023 categories by reported severity
OWASP API Security Top 10:2023 categories by reported severity

The broader OWASP Top 10:2025 view shows Broken Access Control at 27%, Insecure Design at 15%, Security Logging and Alerting Failures at 13%, and Authentication Failures at 11%. The chart separates severity within each category. Every cell is a share of all security-classified findings, not a percentage calculated only within that row.

OWASP Top 10:2025 categories by reported severity
OWASP Top 10:2025 categories by reported severity

We assigned a primary category in each framework. These are analyst judgments from the reported mechanism, not an official crosswalk between the frameworks. Categories organize review questions. Their frequency does not by itself measure business impact, attack feasibility or remediation effort. A category with no assigned findings is not evidence that its risks were exhaustively tested.

CWE classes make the mechanisms more concrete

Common Weakness Enumeration, or CWE, describes classes of software weakness. The leading primary assignments were:

  • CWE-201, sensitive information in sent data: 7%. Information crosses an output boundary that should have withheld it.
  • CWE-862, missing authorization: 7%. A trusted operation lacks the required permission decision.
  • CWE-807, trusting untrusted input in a security decision: 7%. A caller-controlled value receives authority it has not earned.
  • CWE-532, sensitive information in logs: 5%. Diagnostic output carries data that should not enter the log.
  • CWE-636, failing open: 5%. A failure leaves a security restriction unenforced.
  • CWE-778, insufficient logging and audit records, including some monitoring gaps: 5% (mapping confidence Medium). Medium is the dominant mapping confidence, not a claim about every assignment.
Leading primary CWE classes among security-classified findings
Leading primary CWE classes among security-classified findings

Equal rounded shares can hide differences in the underlying figures. These are weakness classes, not descriptions of individual unresolved instances. Secondary CWE assignments do not inflate the shares. An unresolved composite classification remains unverified rather than receiving a guessed CWE number. It remains among all security findings, so uncertainty is not removed to improve the apparent coverage.

Generic patterns to bring into a review

The following examples are invented teaching sketches. They are not product code, exploit instructions or complete fixes. They express safer requirements for familiar failure patterns. They do not identify an open finding or explain what a particular model was thinking.

Missing authorization, CWE-862. A sensitive operation lacks a permission decision. State which identities may perform it and which must be denied. Test both allowed and denied cases. This is a description of the weakness class, not a description of an unresolved instance.

Fail-open behavior, CWE-636. A failure must not silently remove a required security restriction. A failed policy read is not successful absence. Preserve an explicit unavailable state and test the intended refusal and recovery behavior. This is a class-level rule, not a description of an unresolved instance.

Before: a policy read fails; processing uses an empty result.
After: a policy read fails; the decision remains unavailable.

Raw text crossing outputs, CWE-209 and CWE-532. An error message or diagnostic record includes arbitrary exception text. Prefer a safe error category for the consumer and a deliberately constrained diagnostic record. Check returned, stored and logged outputs separately. A correction at the visible response does not establish that every destination is safe.

Work before limits, CWE-770. A program processes all input and trims the result afterward. That limits the response, not the resources consumed. Apply a bound before expensive work, and bound the reader feeding it. Test the accepted boundary as well as oversized input so the protection does not simply reject legitimate work.

Ambiguous input, CWE-20. A permissive conversion accepts more shapes than the contract intended. State accepted types and ranges before processing. This broad CWE is a teaching label here, not a reassignment of every input-related finding. MITRE discourages CWE-20 for real mappings. A more specific mechanism should retain its more specific classification.

Claim and evidence drift is also worth reviewing, but it is not automatically a vulnerability. A summary saying “passed” without a measured result needs correction. That alone does not establish an attacker-controlled security impact. Keeping this distinction is part of making the security results useful rather than dramatic.

Remediation: measured closure, not a safety declaration

As of October 4, 2026, 82% of security findings had merged scoped fixes. Another 12% remained open, including fixes not yet merged. Accepted dispositions accounted for 3% and unknown status for 2%. Those last groups are not counted as fixed. Open findings are shown only as shares of groups with more than ten findings, never as single items. Smaller categories are pooled in the remediation chart.

Remediation shares by OWASP Top 10:2025 category, with smaller groups pooled
Remediation shares by OWASP Top 10:2025 category, with smaller groups pooled

Among security findings with merged fixes, 8% were recorded as partly closed and then closed. In those cases, the initial repair did not finish the recorded work. This supports rechecking the original mechanism after a repair. It does not establish how often all first repair attempts fail, because unresolved attempts are outside the group with merged fixes.

The median recorded finding-to-final-merge interval fell on the same calendar day. The same-day share was 68% among the merged security findings with usable dates. These are date proxies, not precise discovery timestamps. They cannot resolve hours or establish an operational response-time promise. Merge timing is not live remediation timing.

Did vulnerabilities escape earlier AI review?

For 37% of security findings, the provenance records contain a cross-model review relationship. Treat this as a recorded claim about review, not proof that the affected code passed that review before merge. For 10%, no claim was found in the bounded source. For 53%, the evidence is unknown or missing.

Recorded review evidence by category, with smaller groups pooled
Recorded review evidence by category, with smaller groups pooled

This counts findings, not reviewed changes. It therefore cannot measure review effectiveness or a miss rate. Current review descriptions can be edited, can refer to later work, and can include the discovery review itself. Uncertain code origins add another limit. The share proven to have escaped an earlier completed review remains undetermined.

Different reviewers found different things

In the broader sampled other-reviewer ledger, 7% of findings had a source-established overlap with a Codex Daybreak Blue finding. Limited non-overlap observations account for 11%; 82% remain unknown. This comparison group includes test and correctness findings. It is not the set of all security findings used above. Unknown overlap is not a demonstrated miss.

Recorded overlap in the broader other-reviewer sample
Recorded overlap in the broader other-reviewer sample

The same-target narratives show complementary findings. An informed verifier added findings after reading the earlier review. Paired initial reviews sometimes produced different positive lists. Other comparisons described a shared mechanism with different scope or consequence. A final recheck revisited an earlier target. These are case histories, not independent trials. Different briefs, prior knowledge and selected targets prevent a fair model ranking.

Methodology and limits

We studied historical records from a codebase over about a month of finding dates. Repeat observations were linked to distinct findings rather than counted as new discoveries. Mapping used finding text, retained confidence labels and separated non-security findings. In a small masked repeat sample drawn from both reviewer groups, about nine in ten assignments matched. Every mismatch was about whether a finding counts as a security weakness at all, which is the boundary behind the 79% share. This was a same-session check, not independent validation.

The main reviewer was Codex Daybreak Blue. Other sampled reviewers included Claude Sonnet, Claude Opus and Codex models. Writer evidence included Claude Code attribution through commit co-author trailers and other recognized AI tools. A missing Codex writer trailer did not establish no contribution. These records do not support a coding-agent or reviewer ranking. Attribution is incomplete, work exposure differs, and the samples were selected rather than random.

What to ask your team to demonstrate

Start with a security contract, not a request to “look for issues.” Ask who may act on an object, where that permission is checked, what happens when required information is unavailable, and whether limits apply before costly work. Ask the reviewer to trace those answers across callers and outputs.

For a regression test, safely restore the exact mistake it is meant to catch. Require a failure, restore the correction, and require success. Include valid cases. For remediation, record whether the fix is proposed, merged, rechecked or deployed instead of using “done” for every stage.

Use different review perspectives, but ask each for evidence. Keep uncertainty in the report. A vulnerability study can guide better review questions without proving that a product is secure or that a model is the best reviewer.

Shares are rounded to whole numbers and may not sum to 100%.

Talk to us

Explore the Platform

Explore Run-True Decision’s Fraud Decision Engine.

View Platform Overview

Related Articles