← Case Studies
Our work ·

Trustworthy AI for the Ocean

Bridging the Science-Policy Divide with IPOSGPT

A RediMinds Case Study — Published in npj Ocean Sustainability (Nature), June 2026

The Challenge

Thousands of peer-reviewed ocean science papers are published each year, yet the insights within them rarely reach the non-expert policymakers and intergovernmental bodies responsible for ocean sustainability decisions. The translation of scientific evidence into actionable policy is a persistent bottleneck with real-world consequences for marine ecosystems and climate commitments.

General-purpose large language models (LLMs), despite their impressive synthesis capabilities, are fundamentally unreliable for high-stakes scientific and policy domains. When queried on ocean science topics, models such as gpt-4o, gemini-2.0-flash, sonnet-4, and sonar-pro produced fabricated citations at rates exceeding 95%, making their use in evidence-based policy workflows untenable.

Policymakers and intergovernmental bodies require not only accurate answers but also auditable, traceable reasoning grounded in verifiable sources. Existing general-purpose LLM references are not reliable, creating a significant trust deficit in AI-assisted decision-making for ocean governance.

Approach

To develop IPOSGPT, RediMinds and Vital Ocean followed a rigorous, multi-stage methodology:

01

Domain Corpus Curation: A specialized corpus of more than 800,000 documents was assembled comprising curated peer-reviewed ocean science literature, intergovernmental scientific reports, and environmental policy documents. Emphasis was placed on source credibility, coverage of policy-relevant topics, and representation of the global ocean science community, including outputs from UNESCO's Intergovernmental Oceanographic Commission (IOC).

02

Domain-Specific Model Training: IPOSGPT was developed as a domain-specific LLM trained exclusively on the curated ocean science corpus. Unlike general-purpose models, IPOSGPT was designed to restrict its responses to verifiable, sourced content, eliminating the mechanism by which hallucinated references arise in retrieval-free generation pipelines. Development of IPOSGPT took place between December 2024 and April 2025.

03

Anti-Hallucination Architecture: The model architecture enforced strict evidence alignment, ensuring that every generated claim could be traced back to a document within the training corpus. This approach directly targeted the citation fabrication problem observed in baseline models, replacing unconstrained generation with grounded synthesis.

04

Comparative Benchmarking Against Baseline LLMs: IPOSGPT was systematically compared against leading general-purpose models including gpt-4o, gemini-2.0-flash, sonnet-4, and sonar-pro across a 25-query benchmark. Evaluation criteria included citation validity, answer quality, and evidence alignment.

05

Real-World Application — Seychelles Circular Economy Assessment: To validate practical utility beyond benchmarking, IPOSGPT was applied to a real-world policy challenge: a circular economy assessment for marine pollution reduction in the Seychelles. This case demonstrated IPOSGPT's ability to accelerate evidence synthesis for sustainability policy in a geographically and ecologically specific context.

IPOSGPT high-level conceptual system architecture, from user query through keyword extraction, source selection, fact-checking, and the foundation LLM
Solution

IPOSGPT represents a purpose-built, domain-specific AI system that bridges the gap between ocean science evidence and the policymakers who need it. By training exclusively on curated peer-reviewed literature, intergovernmental reports, and environmental policy documents, IPOSGPT delivers synthesis that is both scientifically grounded and policy-relevant.

The system eliminates hallucination at the source rather than through post-hoc filtering, with an architecture that constrains generation to verifiable, corpus-grounded content. This makes IPOSGPT fundamentally different from general-purpose LLMs that repurpose broad internet-trained knowledge for specialized queries, inevitably producing unreliable citations and unverifiable claims.

Developed through an international collaboration spanning RediMinds, the International Platform for Ocean Sustainability (IPOS), Vital Ocean, OSF/CNRS Foundation, IOC-UNESCO, Sorbonne AI, and other leading institutions, and supported by the AXA Foundation, CNRS Foundation, and French Ministry of Research, IPOSGPT represents a model for how domain-specific AI can be responsibly deployed in high-stakes, evidence-dependent governance contexts.

IPOSGPT product interface, showing the query screen and source-filtering settings used by policy researchers
Results
  • Zero fabricated references: IPOSGPT produced zero hallucinated citations across tested policy-relevant queries, compared to a rate exceeding 95% invalid citations among all four baseline general-purpose models (gpt-4o, gemini-2.0-flash, sonnet-4, sonar-pro).
  • Second-ranked answer quality: Despite its strict evidence alignment constraints, IPOSGPT achieved second-ranked answer quality across the 25-query benchmark, demonstrating that scientific rigor and answer quality are not in conflict.
  • Real-world policy acceleration: The Seychelles case study confirmed that IPOSGPT can materially accelerate circular economy assessment workflows for marine pollution reduction, enabling faster, evidence-backed policy development in resource-constrained environments.
  • Validated trust framework: The results collectively demonstrate that domain-specific LLMs can deliver trusted, accountable synthesis for high-stakes sustainability policy, directly addressing the core challenges of hallucination, source integrity, and decision accountability.
Conclusion

The publication of this research in npj Ocean Sustainability (Nature) marks a significant milestone in RediMinds' mission to bring rigorous, trustworthy AI to high-stakes decision-making domains. IPOSGPT demonstrates that the hallucination and reliability challenges that have limited LLM adoption in scientific and policy contexts are solvable, not through general-purpose scaling, but through domain-specific design, curated knowledge, and evidence-grounded architecture.

As ocean sustainability moves higher on the global policy agenda, the need for AI systems that policymakers, scientists, and intergovernmental bodies can genuinely trust will only intensify. IPOSGPT offers a replicable model for deploying trustworthy AI at the science-policy interface, one that prioritizes accountability, traceability, and scientific integrity at every layer of the system.

Reference

Pankajakshan, A., Danielson, J., Wang, Z., Singh, P., Mehra, M., Pawar, B., Reddiboina, M., De Buyser, E., Tawil, R., Mehta, S., Barnhill, K. A., de Lisle, M., Gaill, F., Ortuno Crespo, G., Enevoldsen, H. O., Fresquet, X., Astier, E., González Estrada, A., & Brodie Rudolph, T. (2026). Trustworthy AI for the ocean: bridging the science-policy divide. npj Ocean Sustainability. https://doi.org/10.1038/s44183-026-00195-0