JATS XML in Scientific Publishing Workflows
The Journal Article Tag Suite (JATS) is the global technical standard for structuring, archiving, and exchanging scientific journal articles in XML format. By establishing a common semantic grammar, JATS ensures that scholarly content remains machine-readable, interchangeable across platforms, and universally accessible. This article explores how JATS simplifies scientific publishing workflows—from manuscript submission and automated typesetting to multi-format rendering, indexing, and long-term digital preservation.
What Is JATS XML?
Maintained by the National Information Standards Organization (NISO) as standard Z39.96, JATS originated from the National Library of Medicine (NLM) DTD. It provides an extensive set of XML tags specifically designed to describe the content and metadata of scholarly articles. A standard JATS XML document is organized into three primary sections:
<front>: Contains critical metadata such as article titles, author names, affiliations, abstracts, funding details, and publication dates.<body>: Structures the narrative content, including sections, paragraphs, figures, tables, formulas, and supplementary materials.<back>: Handles supplementary references, appendices, acknowledgments, and author notes.
Establishing a “Single Source of Truth”
Before JATS, publishers relied on proprietary formats or fragmented pipelines, creating bottlenecks when converting manuscripts into multiple output types. JATS XML functions as a single source of truth within a publishing workflow.
Once an accepted manuscript is marked up in JATS XML, automated typesetting engines can transform the same underlying data into various distribution formats, including PDF for print/download, responsive HTML for web reading, and EPUB for mobile devices. Any editorial corrections applied to the XML master document automatically propagate across all derivative formats.
Ensuring Interoperability Across Systems
Scientific publishing involves a complex network of stakeholders, including manuscript submission platforms, editorial boards, external typesetting vendors, institutional repositories, and indexing services. JATS standardizes the data exchange between these disparate systems.
Because platforms like PubMed Central (PMC), Crossref, and various academic search engines support JATS, publishers can transmit metadata and full-text files without custom mapping or data translation layers. This interoperability eliminates manual data entry, speeds up publication timelines, and reduces human error during production handoffs.
Enhancing Semantic Discoverability and Machine Readability
Modern research requires automated discovery and computational analysis. JATS enriches research papers with semantic meaning by tagging specific entities—such as chemical formulas, mathematical notations (via MathML), digital object identifiers (DOIs), and ORCID identifiers for researchers.
This deep tagging enables search engines and academic aggregators to accurately index content, extract citation graphs, and facilitate text and data mining (TDM). Reference linking becomes fully automated, allowing readers and algorithms to instantly navigate from in-text citations to original source materials.
Future-Proofing and Digital Preservation
Digital preservation formats must remain accessible independent of proprietary software. Because JATS XML is an open, human-readable, plain-text standard, it guarantees that scholarly literature can be archived and retrieved decades into the future without data degradation. Major global archives, including Portico, PMC, and national libraries, utilize JATS as their core ingestion format to safeguard the scientific record.