Skip to content

Written by Mark Collins, Director of Academic, Virtusales 17 August 2026

Metadata may not be the flashiest topic. In over 25 years of working with publishers on their systems, I've yet to meet someone who describes it as thrilling — but its importance is hard to overstate. I've had plenty of conversations where it turned out to be the thing sitting quietly behind a problem they couldn't quite pin down.

We've just published a white paper that tries to make sense of why that keeps happening. It's called Metadata as Infrastructure: Powering Governance and Growth in the AI Era, and it's free to download from our website. This post is my attempt to summarise the conversation, though I'd encourage you to read the full version if any of it sounds familiar.

The short version is that four things publishers are worrying about right now – ensuring their trusted research gets discovered, protecting their rights in an AI world, meeting their legal obligations on accessibility, and just the day-to-day inefficiency of running a publishing operation – all have metadata somewhere in the middle of them. Not as the whole answer, but as the bit that keeps making the other answers harder.

Discoverability

Let’s start with discoverability, because it’s the one people are most familiar with. Nielsen BookData looked at the top 100,000 bestselling ISBNs in the UK and found that titles with complete metadata sold roughly twice as many copies as those without; a cover image alone made about a 94% difference to average sales. For university presses and learned society publishers – where a researcher might be hunting for one very specific title through a library catalogue, a discovery platform, or something like JSTOR – being hard to find is a real commercial problem, regardless of how good the content itself is. I don’t think I need to make that case at length to this particular audience.

Rights and AI

The AI side of things is a different kind of conversation and, honestly, a harder one. I’m not a lawyer. I don’t know how the various court cases will land. What I do find myself saying to publishers, though, is that the legal question and the metadata question are tangled together in a way people don’t always see at first.

Say you’ve decided you don’t want your content used for AI training, or that you’re fine with it but only on certain terms. That position has to be readable by a machine, or it doesn’t do anything. An opt-out sitting in a contract somewhere won’t stop a crawler. Think of it like a ‘no cold callers’ sign written in pencil on the inside of your front door. You’ve technically put up a sign. But there’s no guarantee that anyone coming to the door can see it.

And then there’s the money side, which I don’t hear discussed enough in these conversations. Wiley made close to $100 million from AI content licensing deals in 2024, and Informa and Taylor & Francis did similar – and whilst you might think those deals were exactly right, or have reservations about them, either way, those publishers could sit at that table at all because their rights data was clean and structured enough to put on it. Publishers whose data isn’t in that shape don’t get to decide. The decision gets made for them.

Accessibility

Something I find myself returning to fairly often at the moment is the European Accessibility Act and, in North America, the Americans with Disabilities Act Title II. Both set real dates. Both require you to actually show your working, not just declare a commitment. The alt text, the structured tagging, the descriptors – it all has to be in the product itself, in a form someone can verify, which puts it squarely in the metadata conversation.

A European survey from 2024 found that only 37% of ebook-producing publishers are actually putting out accessible ebooks yet – which surprised me when I saw it, because awareness of the requirement is much higher than that – and fewer than one in ten published works globally carries any accessibility metadata at all (things like image descriptions, font information and so on). Going back through years of content to add all that retrospectively is a big job. Much more sensible to make it part of normal workflow from here on. The publishers I know who’ve done that are considerably less anxious about the deadlines than those who haven’t. 

And there's a side benefit that doesn't always get mentioned: a lot of the metadata work you do for accessibility – the descriptions, the structured tagging, the additional detail about a product – also helps with discoverability. The same record that becomes more accessible to a visually impaired reader becomes more discoverable to an algorithm too. So it's not always as costly a project as it first appears.

Workflow and the hidden cost of manual processes

Which brings me back to the spreadsheet. When I visit publishers and ask how a process works, and someone mentions a spreadsheet, I’ve got into the habit of asking a follow-up: what happens when the person who looks after that spreadsheet moves on? The answer is usually an uncomfortable pause. Sometimes a nervous laugh.

The white paper puts some rough numbers on this. A mid-sized publisher with around 3,000 titles and 15 distribution channels is probably losing somewhere in the region of £193,000 a year through missed launch windows, and spending around 600 staff hours on sorting out metadata errors, chasing feed rejections, and reconciling things that should already match. Kogan Page cut 250 of those hours just by automating how their ONIX data goes out to external systems – that’s one process, at one publisher. The hours that frees up are hours those people can spend on work that actually needs them.

The reason I keep coming back to metadata as the common thread is that it genuinely is. Discoverability, rights, legal compliance, workflow efficiency – they all need the same records to be right. Fix the records properly and you’re not running four separate improvement projects. You’re doing one thing that helps across all four. 

And I'd add a fifth, which doesn't come up often enough: trust. Misinformation is a real and growing problem, and publishers – scholarly and academic publishers especially – are actually in a strong position to be a reliable signal against it. But that only works if the metadata travels with the content and says something meaningful about where it came from. Quality metadata is how trusted content gets found by the right people, rather than getting lost while something less reliable surfaces instead. 

You can download the full white paper for free at virtusales.com. It covers all four areas in a lot more detail, with the data behind it — and if you work in publishing, it's probably worth a look. 

I'll also be at the ALPSP Annual Conference in Manchester in September, taking part in a session on the Thursday morning — AI-Ready Books: Practical Metadata, Rights, and AI Strategies for Impact — which picks up a lot of what's in this post. If you're coming to the conference, do come along, or come and find me there if you'd like to chat in-person. 

If you're not going to be in Manchester, I'd be happy to have a chat. 

Contact me.


About the author

Mark Collins is Director of Academic at Virtusales and has over 25 years of publishing industry experience. He has implemented Virtusales’ BiblioSuite software for global publishers including Elsevier, Macmillan and Bloomsbury, as well as scholarly and academic leaders like Oxford University Press, Harvard University Press and the MIT Press. Mark previously held various roles at Wiley, where he focused on implementing and delivering global technology solutions across the business. He is also an ALPSP and SSP mentor.