The short version
The question I keep getting about Coffeetable is why it doesn’t already exist. Google scanned the world’s libraries starting in 2004. Amazon was searching inside 120,000 books in 2003. So why is the thing that puts pages from four authors in front of you a plugin built by one person in 2026?
The honest answer is that it did exist. Google built the hard part twice, won the lawsuit that made it legal, and then switched it off.
Almost everything Coffeetable does mechanically was possible by 2010. Scanning, OCR, a full-text index over millions of pages, passage ranking, page-level similarity across scanned books. Semantic search is not an LLM invention either. Latent semantic indexing was published in 1990 to solve exactly the problem where the words in your question don’t appear in the passage that answers it.
What was missing was permission, and the shape of the permission problem is stranger than I expected before I went and read the case files. It isn’t that publishers said no to book search. They said yes to book search, repeatedly, when it ended at a checkout. It isn’t that a court ruled the idea illegal either. A court ruled you can’t get it by default.
And the part that reframed the whole thing for me: Google’s fair-use victory rests on Google Books not being good enough to read. The Second Circuit’s reasoning leans on the restrictions — fragments of a page, snippets cut by machine rather than chosen, one page in every ten permanently withheld so you can’t reassemble the book by repeating queries.
Google spent a decade and a fortune establishing the right to not show you the book.
That holding is the ceiling over every book product built since, including mine. Below is the exact chronology, because with this story the dates are the argument.
The exact Google story
2004: copy first, ask later
In December 2004 Google announced library partnerships with Michigan, Harvard, Stanford, Oxford and the New York Public Library. The libraries handed over books. Google kept the scans and the machine-readable text, and gave the library a digital copy back.
For books that publishers submitted themselves through the Partner Program, the publisher decided what a reader could see. For library books, Google scanned in-copyright works without asking each author or publisher first. That sentence is the entire lawsuit.
The Authors Guild sued in September 2005. Five publishers, through the Association of American Publishers, sued in October. Neither case was really an argument about whether search is useful. They were arguments about who sets the price of a copy, and whether a technology company gets to build a market and make rightsholders opt out of it afterwards.
2008: the settlement was this product
In October 2008 the parties proposed a $125 million class settlement. I expected a payoff for past scanning with some indexing rules attached. It was not that. It was a blueprint for a licensed digital reading market, and it was specific:
- a Book Rights Registry to identify rightsholders and pay them;
- at least $60 for each work already scanned;
- consumers buying digital access to individual books;
- institutional subscriptions;
- free access terminals in public libraries;
- previews of up to 20% of a book;
- 63% of the revenue to rightsholders;
- and room for later models including subscriptions, print-on-demand, custom publishing, downloads, summaries and compilations.
So the market wasn’t unimagined. It was drafted, priced and agreed by the people who owned the books.
2011: a judge kills it on procedure
In March 2011 Judge Denny Chin rejected the amended settlement. Not because digital reading was bad for readers. Because a lawsuit about past scanning was being used to hand over sweeping rights to future uses, including the works of authors who were not in the room. Orphan works. Foreign rightsholders. Opt-out rather than opt-in. Chin said as much: an opt-in structure would answer most of the objections.
That distinction is the one useful thing anyone building today inherits from 2011. No court has held that a licensed reading layer over books is unlawful. A court held that you cannot acquire one by default.
2012: publishers take their control back
Google and the publisher plaintiffs settled separately in October 2012 on confidential terms. Publishers got to decide, title by title, whether a scanned book stayed in Google Books, came out, or went further into preview and sale. Control was restored. The universal market was not built. The Authors Guild kept litigating.
2015: Google wins by staying small
The district court found fair use in 2013. The Second Circuit affirmed in October 2015. The Supreme Court declined the case in April 2016. Google won, decisively, and it is worth being precise about what it won on.
Making complete internal copies was fair use for full-text search and for the snippet view as then designed. The court’s reasoning rests on the design limits: the public never receives the scans; a result shows a fragment of a page; snippets are divided mechanically instead of selected as complete ideas; one snippet on every page and one page in every ten are permanently withheld; a reader cannot iterate queries to reconstruct the book. The display tells you whether a book is relevant without standing in for reading it. The court was explicit that the snippets give enough context to evaluate interest without revealing enough to threaten the author’s copyright interest.
2018: Google ships the conversation
In April 2018 Google Research launched Talk to Books. You typed a question or a statement in ordinary language. A neural language model compared it against every sentence in more than 100,000 books and returned the passages that read like answers, with surrounding context and a link to the book. No keywords required.
That is the retrieval half of Coffeetable, eight years early, at Google scale. By Google’s own account, millions of people used it, the semantic matching worked, and the underlying techniques went on to ship inside Shopping, Maps, Gmail and Hangouts.
2023: Google turns it off
Talk to Books closed in June 2023. Google’s stated reason is that the experiment had taught them what it had to teach. Google does not say publisher opposition or bad economics killed it, and I’m not going to invent that. What the public record supports is narrower and more useful to me: the interaction worked at scale, and Google treated it as research rather than as a book business.
So the accurate version of “nobody built this” is that Google built the retrieval, established the law, helped draft the market, and then decided it did not want to be the rights clerk for every page it could serve. Not one of those steps failed on technology.
The line the court drew
Reading a book into a machine and showing a human words out of it are two different rights. Every outcome in this area since 2014 turns on that split, and once you see it the cases stop looking contradictory.
In 2014 the Second Circuit held that HathiTrust’s full-text searchable database of scanned library books was fair use. An ordinary user could learn which books and which pages contained a term and received no text at all. Full-text access was permitted for certified print-disabled readers as a separate purpose with its own analysis.
Ten years later the same court held that the Internet Archive’s controlled digital lending was not fair use, even under one-owned-to-one-loaned. The digital copies served the same purpose as the originals, reproduced whole books, and competed with the library ebook licensing market that already exists. The Archive dropped its appeal in December 2024.
There is a second rule I have to design around, which is that purpose does not travel. In July 2026 Hachette, Cengage, Elsevier and the author Scott Turow filed a proposed class action alleging that Google used works it obtained for Google Books and Play Books to train Gemini. Those are allegations, not findings. But the position behind them is unmistakable, and I think it’s correct: permission granted for one product does not quietly extend to the next one.
The Anthropic litigation is sometimes offered as cover here, and it isn’t. A $1.5 billion settlement over roughly 500,000 works is not a licence for anything I would do. Training a model to produce different words is a different act from retrieving the original words and putting them on a screen.
Everyone else who tried
Nine attempts, and they sort themselves cleanly by what the money did rather than by how good the technology was.
| Attempt | How it was paid for | What happened |
|---|---|---|
| Amazon Search Inside the Book 2003 | Publisher opt-in inside a store; the preview protects the sale | Alive. Launched with 190 publishers, and Amazon reported that sales of included titles grew 9% faster than the control group in five days. Discovery is easy to approve when it ends at a checkout. |
| Google Books 2004 | Fair use for search, permission for previews | Alive as discovery infrastructure. What it offers a publisher is control and reporting, with the reader sent out to a seller. |
| Microsoft Live Search Books 2006–2008 | Library and publisher programmes, funded as search investment | Closed. Microsoft said it needed a sustainable model for the search engine and the consumer and the content partner at once, and moved to verticals with higher commercial intent. |
| Small Demons 2011–2013 | Publisher partnerships, hoped-for affiliate revenue | Closed after an acquisition fell through, having raised more than $2 million. Gorgeous literary metadata, no revenue of its own. |
| BookLamp 2007–2014 | Text analysis sold B2B | Acquired by Apple. Recommendation derived from text is a store feature, not a destination. |
| Kindle samples and X-Ray 2011 | Rights already inside Kindle distribution | Alive. Navigation is valuable after the purchase, sampling right before it. |
| Safari / O’Reilly 2001 | Licensed corpus, reader pays a subscription | Alive. Search and read across many publishers works when the corpus is licensed and someone is paying. Easiest in professional categories. |
| Talk to Books 2018–2023 | Research budget | Closed. Millions of users, no business attached. |
| Internet Archive lending | Fair use instead of ebook licences | Lost in court. A nonprofit motive and a one-at-a-time limit did not save it. |
Everything above settles into one of five stable positions. Index it and never let the index become the book, which is Google and HathiTrust. Show a controlled sample to cause a purchase, which is Amazon and Google’s partner previews. Show everything because somebody paid for licensed access, which is Kindle and O’Reilly and every library ebook platform. Use only public-domain or openly licensed work, which is Project Gutenberg and Coffeetable today. Or show it without permission and stay legally unstable, which is the shadow libraries and the lending theory that just lost.
There is no sixth example. Nobody has sustainably assembled arbitrary verbatim pages from in-copyright trade books into something a reader can take away, without a licence or a purchase sitting behind it.
What the LLM actually changed
Three things, and they’re real.
Intent. “I have to fire a friend on Monday” becomes literary concepts, candidate works and plausible search language without the reader knowing a single keyword. That was the hardest unsolved part of every earlier system and it is now nearly free.
Editorial judgment. A model can read what came back and put four passages in an order that argues something. Talk to Books returned ranked sentences. Coffeetable returns a sequence with movement in it.
The interface. You ask in your own words and you push back in your own words, which is the difference between a database and a librarian.
Now the part that gets skipped. None of those three produces the right to ingest, retrieve, display or hand over the words. The 2015 opinion reads exactly the same in 2026. If anything the model makes my legal position worse rather than better, because it makes the output better, and better here means closer to reading.
Where Coffeetable sits on that line
What it does today: takes a reader’s actual situation, picks real books, searches their full text, pulls exact page labels, chooses the passages that answer the question, and binds verbatim pages from two to four books into one short reading you can finish in the conversation. The public catalogue is hard-limited to reviewed public-domain Gutenberg editions.
Inside “search books” there are six different jobs, and they are not equally negotiable:
- Recommend. “A dark novel with beautiful prose, no fantasy.” Titles and reasons come back. Nobody objects to this.
- Navigate. “Where does this book talk about envy?” Locations come back. Approved for twenty years.
- Sample. “Let me judge the writing.” A controlled page or chapter. Familiar, and already licensed everywhere.
- Recover. “Find the scene where the teacher explains the girl’s illness.” The actual pages. Acceptability depends entirely on how many.
- Synthesise. “Make me a reading on grief from four authors.” A new sequence of verbatim pages. This is the part readers love.
- Carry it away. “Put that on my Kindle.” A durable file that leaves my controls behind.
I’m not going to argue that steps 5 and 6 are discovery. Presenting a cross-book compilation as a preview is how you lose a rights department in four minutes, and they’d be right.
What a publisher would actually have to sign
The way through is not a cleverer fair-use theory. It’s a licence narrow enough that somebody can sign it and useful enough to be worth testing. Four things have to be true.
Two switches, enforced by the renderer. The system has to be able to say “this book answers your question” without assuming it may show you the book. Every title carries its own rules: who owns it, which edition, which territories, whether it can be indexed, whether it can be embedded for retrieval, how many words and pages can ever be visible to one reader, whether it can appear alongside other titles, whether pages can be exported, where the buy link points, when the licence expires and the text is deleted. The rendering service enforces those. Never the prompt. A limit a model can be talked out of is not a limit.
A portfolio, not a permissions queue. Ask one rightsholder for 100 to 500 backlist titles, one territory, ninety days, a hard page cap per title per reader, no raw-text access, no training, revocation whenever they want it. Backlist because it is the part nobody is merchandising this quarter, which makes incremental demand a believable claim instead of a hope.
Export is its own licence. An expiring reader and a downloadable EPUB are not the same object. The file outlives my controls and starts to resemble a custom publication. For a first pilot the honest version sends a purchase link to the device, or builds the file only from public-domain pages.
Somebody has to answer the substitution question. This is the actual reason limits stay conservative. A publisher can see a buy-link click. A publisher cannot see whether the reader who got six excellent pages from four books bought all four, bought one, or went to sleep satisfied and bought none. Nobody has measured it, and it is measurable: run titles only, titles plus a tiny snippet, and titles plus a licensed page-first reading, then report per title what happened. If the page treatment doesn’t beat the control on something a publisher already cares about, I don’t have a business, and that is a cheap thing to find out early.
None of this is a fantasy about publishers loving AI. The industry’s answer to AI has never been a flat no. Wiley booked $40 million of AI content licensing in fiscal 2025, up from $23 million the year before. HarperCollins went to individual authors for a limited AI licence, and the Authors Guild used that episode to argue the rights often sit with the author rather than the publisher. The Association of American Publishers told the White House in 2025 that voluntary licensing markets for publisher content already exist and should be encouraged. Kindle Unlimited already pays out by pages first read, so per-page accounting is not exotic to anyone’s royalty system.
Publishers license when the grant is explicit, the files are controlled and the value is measurable. Nothing on that list is impossible for a small product. All of it is slow.
And talking to a publisher is not the same as talking to the owner. Older contracts frequently didn’t convey electronic rights at all, and the RosettaBooks case established that a grant to publish “in book form” didn’t stop authors licensing ebooks separately. Penguin Random House’s own permissions process treats an anthology as its own category of use, can require one request per book, and says plainly that some uses can’t be cleared without the author. A cross-book reading looks less like a preview and more like a custom digital anthology. I’d rather put that sentence in front of a rights department early than hide it behind the word AI and get found out.
What I’m doing about it
Coffeetable is public-domain only right now, and that is a decision rather than a placeholder for a licence I’ve already got. Gutenberg is north of 70,000 books, and a lot of what people actually bring to Coffeetable — grief, ambition, discipline, love, how to think about a decision — was written before 1930 by people who wrote about it better than the current bestseller list does.
What I want next isn’t more retrieval code or a nicer reader. It’s two facts I don’t have. Whether readers come back on their own. And whether one rightsholder will put a real title list in writing.
If someone grants 100 commercially relevant titles on terms that leave the product feeling like itself, and readers return without being prompted, the thesis holds and it’s worth everything that follows. If readers return but every rights conversation collapses into per-title and per-author clearance, then the public-domain version is a genuinely good product and a smaller company, and I’d rather run that honestly than keep implying the universal library is one deal away. If readers don’t come back, none of this legal history matters and I should stop.
The one road I won’t take is the one that needs a novel fair-use theory for consumptive display. Google spent ten years and an enormous amount of money establishing a narrower right than that, and I have neither.
The part that stings
“Nobody built it” is the comfortable story. Here is the real one.
Everything mechanically necessary existed by 2010. The market for it was drafted in 2008 by the lawyers of the authors and publishers who were suing Google. It died in 2011 over a procedural question about people who weren’t in the courtroom. The conversational version shipped in 2018, reached millions of readers, and was switched off in 2023 because it wasn’t strategic for the company that happened to own it. And the case that made any of it legal was won on the argument that the product deliberately stops short of being worth reading.
That is not a technology gap. It’s a permission gap with a twenty-year head start.
Which is worse news in one way, because I cannot out-engineer it. It’s better news in the way that matters: the gap isn’t held shut by a moat. It’s held shut because everybody with the rights had no reason to try, and everybody with a reason had no rights. That is the kind of gap one person with a small catalogue and a specific promise can push on.
Whether it opens, I don’t know yet. Going to find out.
Coffeetable runs inside ChatGPT and Claude, on public-domain books. Add it in a minute, or read why it never summarizes.

