Stockholm-based legal technology startup Legora has demonstrated a multi-document financial audit workflow using OpenAI's newly discussed GPT-6 Astra model, successfully completing a financial statement tie-out across 41 documents in a matter of minutes. The exercise, documented by OpenAI News and reported via StartupHub.ai, tested the agent's capability to ingest complex supporting schedules, trial balances, and prior-year accounts simultaneously without suffering from the context degradation that typically limits legal and financial automation tools.
Financial statement tie-outs are traditionally intensive exercises that require compliance teams and legal engineers to spend entire evenings or multiple days manually checking every figure in draft accounts against supporting documentation. According to Legora legal engineer Percevale Perks, the task is grunt work that survives because mistakes carry severe financial penalties.
Benchmarking Agentic Reasoning at Scale
To evaluate the capabilities of GPT-6 Astra, Legora subjected the model to its internal Benchmark for Agentic Reasoning, known as the BAR suite. This benchmark measures end-to-end performance across real legal workflows rather than relying on isolated question-and-answer interactions. On the 41-document tie-out workflow, Astra scored nearly 40 percent higher than the previous model generation.
Across the entirety of the BAR benchmark suite, however, the average performance lift was more modest at approximately 3 percent. In a controlled test involving four planted discrepancies—including a notable 500,000 British pounds revenue note gap—Astra successfully identified all four errors while retaining every correct check from the prior model and adding roughly 50 additional correct checks in a single pass.
The Challenge of Unstructured Real-World Documents
While the benchmark results highlight significant advancements in multi-file context handling and structured output generation, industry observers note that synthetic test environments do not fully replicate messy, year-end corporate closes. Real-world financial files frequently feature inconsistent naming conventions, fragmented formatting, and scanned PDFs rather than clean digital exports.
For enterprise developers, the core engineering signal centers on balancing expansive context windows with structured, auditable outputs. Auditors require transparent logs that can be independently verified, making the reliability of data ingestion as crucial as the underlying model's reasoning capabilities.
Human Supervision and Enterprise Expansion
Legora maintains that the AI agent is designed to execute exhaustive comparisons while human professionals retain final decision-making authority. This human-in-the-loop framework forms the foundation of Legora's broader expansion beyond contract review into tax, compliance, risk management, and audit operations.
Positioned as an agentic operating system rather than a single-purpose tool, Legora reports usage by more than 100,000 professionals across 1,800 legal departments in over 50 markets. The platform's architectural approach includes integrations into enterprise software stacks, ranging from Intapp Walls to Google Cloud's Gemini Enterprise for Legal.
Security Implications and Future Outlook
Deployment of advanced models like Astra also brings heightened exposure. The demonstration follows recent disclosures from OpenAI flagging critical cyber risks associated with the Astra model, highlighting the security challenges inherent in granting autonomous agents broader access to enterprise systems.
As organizations weigh the efficiencies of automating initial financial passes against security exposure, the central test remains economic. The industry must determine whether enterprise clients will continue paying for rigorous audit trails and human-reviewed verification layers as foundational AI models advance.