What Is the NCBI JATS XML Standard for Journals?
The Journal Article Tag Suite (JATS) XML standard, initially developed by the National Center for Biotechnology Information (NCBI) and the National Library of Medicine (NLM), is the global standard for structuring, archiving, and exchanging digital scientific and scholarly literature. By providing a common technical framework, JATS XML enables publishers, libraries, and databases to store journal article content and metadata in a consistent, machine-readable format that ensures long-term preservation and seamless interoperability.
The Core Purpose of JATS XML
The primary purpose of JATS XML is to solve the fragmentation of digital publishing formats by defining a single, standardized markup language for journal articles. As scholarly communication moved online, the industry required a universal format to ensure articles could be preserved, discovered, and shared across diverse technical systems.
1. Universal Interoperability and Data Exchange
Prior to JATS, publishers used proprietary formats, making it difficult to transfer content between organizations. JATS provides a shared schema recognized worldwide. This uniformity allows publishers to seamlessly submit articles to major repositories and indexing services, such as PubMed Central (PMC), Europe PMC, and institutional archives, without needing custom data transformations for each destination.
2. Long-Term Digital Preservation
Proprietary formats and design-heavy files like PDFs can become obsolete or difficult to parse as technology evolves. JATS XML separates the content and structure of an article from its visual presentation. Storing articles as plain-text XML based on the NISO Z39.96 standard ensures that scholarly records remain accessible, parsable, and intact indefinitely.
3. Rich Metadata and Machine-Readability
JATS enables granular semantic tagging of every element within an article. Rather than merely displaying text, JATS identifies specific data points, including: * Author names, ORCID identifiers, and institutional affiliations * Funding information and grant numbers * Detailed citation structures and persistent identifiers (DOIs) * Abstracts, keywords, figures, tables, and licensing terms
This granular tagging makes content fully machine-readable, empowering automated research, text-and-data mining, and enhanced indexing by academic search engines.
4. Efficient Multi-Format Publishing
JATS XML serves as the single source of truth in modern publishing workflows. From a single validated JATS XML file, publishers can automatically generate multiple end-user formats, including web-ready HTML, print-ready PDF, EPUB for mobile devices, and plain text for indexing. This streamlines production and reduces formatting errors across platforms.
5. Enhanced Discoverability and Accessibility
Because JATS XML clearly defines article components, search engines and discovery platforms can index scholarly works with high precision. Furthermore, structured XML directly supports web accessibility requirements, enabling screen readers and assistive technologies to interpret complex scientific structures, formulas, and data tables accurately.