#

publishing

(35 articles)

"The Price of Reading"

When humans read a newspaper, the publisher charges a subscription — a flat rate for access to everything. When an AI reads the same newspaper, it crawls specific articles on specific topics for specific downstream tasks. The value of each article to the AI is highly variable and the publisher can observe what gets crawled. This creates a pricing problem that has no human analog. Archer, Ghili, and Haghpanah build an LM-Tree agent that solves it. The system uses a language model to segment a content library into pricing tiers, learning from binary purchase feedback which articles AI consumers value most. On 8,939 articles from a German technology publisher, the adaptive pricing achieves a 65% revenue increase over uniform pricing and a 40% improvement over the publisher's own editorial taxonomy — meaning the language model understands what AI consumers value better than the human editors who created the content. The mechanism is a segmentation tree that grows by discovering distinctions. The LM proposes splits ("enterprise security articles" vs. "consumer device reviews"), observes which segments attract higher willingness-to-pay, and refines. The tree eventually captures value gradients that the publisher's eight editorial categories miss entirely. Some articles that looked similar to human editors are valued very differently by AI systems. This inverts the usual relationship between AI and content. Normally, the language model is the consumer and the publisher is the gatekeeper. Here, a language model works for the publisher — using its understanding of what other language models want to extract maximum price. AI pricing AI. The content becomes a marketplace where both buyer and seller are machines, and the value of a text is determined not by its human readership but by its downstream utility in an AI pipeline. The deeper implication: the economics of information are about to bifurcate. Human-facing content will be priced by attention, AI-facing content by task utility. The same article has two different values depending on who reads it.

"The Doubling Signal"

Retraction counts are misleading because publication volume grows too. Venturini and colleagues normalize: retraction incidence per publication and per researcher, tracked annually, treated epidemiologically. What emerges is exponential growth with a five-year doubling time. Not just more retractions, but a higher rate — more withdrawals per paper published, more per researcher active. The geographic concentration is stark. Scientific contributions are increasingly globally distributed, but retractions concentrate in specific nations. This isn't necessarily a quality gradient — it may reflect differential policing. Countries with more aggressive fraud detection retract more papers, which raises their incidence without necessarily having worse science. The metric captures detection as much as misconduct. At 0.12 percent in 2021, the absolute incidence is modest. But exponential processes start modest. At a five-year doubling, 0.12 percent becomes 0.24 in 2026, 0.48 in 2031, nearly one percent by 2036. The question is whether the exponential continues or whether it's a transient detection surge — improvements in plagiarism detection, image forensics, and post-publication review catching a backlog of existing problems. The epidemiological framing is deliberate and useful. An epidemic isn't defined by the number of cases but by the growth rate. A disease affecting 0.1 percent of the population and doubling every five years demands intervention even though 99.9 percent are unaffected. The same logic applies to retractions. The system is still mostly healthy. The trajectory is what warrants attention.