The Cat Saw a Dog
In March 2013, in a converted chocolate factory in Moscow, I put a slide up that said we cannot tell whether the cat saw a dog or a dog saw a cat — and told the room that word order was the future of text mining. I was right about the direction and wrong about the mechanism. The more useful story is what I had already learned, by then, about which ideas actually get published.
Every researcher has a moment where they said something in public that later looked either prescient or ridiculous, and usually cannot remember which. I have one, and unusually there is a paper trail, so I can check.
This is the story of how I got into this field, the disillusionment that arrived roughly a year into the PhD, and a claim I made out loud in Moscow in 2013 that I still think about. I am going to be careful about the part where I appear to have been right, because that is exactly the sort of story people tell themselves badly.
A thesis that was not about mathematics
I was an undergraduate at Sikkim Manipal Institute of Technology, in my penultimate year, when I read Brian Pinkerton’s work on WebCrawler. Not a textbook, not a course — a dissertation, downloaded and read the way you read something nobody assigned you.
WebCrawler was the first comprehensive full-text search engine for the Web. Pinkerton started it at the University of Washington in 1994 and presented it that October at the Second International WWW Conference in Chicago.1 The dissertation came six years later, after the thing had become a commercial product, and its preface says something I did not appreciate at the time and have thought about ever since:
He is apologising, gently, for having built something that worked instead of proving something that was elegant. He even lists solutions that were theoretically promising and never deployed. I read that as an invitation. Then I read the PageRank work — the Stanford technical report, and Brin and Page’s description of the Google prototype, which notes almost in passing that “very little academic research has been done” on large-scale search engines.34 Two of the most consequential systems in computing, and both papers open by observing that the academy was not especially interested.
So I built one. Between 2007 and 2009, on a research placement in the R&D division of Tata Steel in Jamshedpur, I wrote a multithreaded web crawler as an indigenous search solution, and a localised search engine that ranked results by where you were asking from. It got written up in the Hindustan Times under a headline I did not choose.5 It was not sophisticated. It worked, and I wanted to spend my life on this.
The Mario problem
The PhD started at the Chinese University of Hong Kong, and the first idea I had that felt genuinely mine came from a video game.
Growing up on Super Mario, every player internalises something without being taught: not all jumps are equal. Two ledges close together and low down are trivial. Move the platform further away, or higher up, and it suddenly demands precision. Distance and height are what make a jump hard.
Reading, I thought, is the same terrain. Some sentences are short hops between everyday words; others ask you to leap a wide conceptual gap or climb onto a ledge of dense technical vocabulary. So model the document as terrain and rank text by how hard the jumps across it are. That became a CIKM 2011 paper — four pages, an unsupervised ranking method based on a technical difficulty terrain — and the seed of everything after it.6
I was, briefly, extremely pleased with myself. And then I learned how this actually works.
The part nobody tells you in the first year
Here is the thing I want to say plainly, because I have watched a decade of PhD students discover it the slow way.
The novelty of an idea and its publishability at a top venue are two different variables, and they are less correlated than you would like. What reliably travels is technical weight: a nonparametric prior, a derivation, a sampler, a proof, a table with enough baselines. An idea that is genuinely new and expressible in a paragraph is a hard sell. An idea that is a modest variation but arrives wearing four pages of mathematics is an easy one.
I am aware this is the sort of claim a bitter person makes about reviewers, so take it from people with data rather than from me.
- Lipton and Steinhardt named the pattern mathiness — borrowing the term from the economist Paul Romer — and defined it as “the use of mathematics that obfuscates or impresses rather than clarifies”. They describe spurious theorems inserted “to lend authoritativeness to empirical results” even where the theorem does not support the paper’s actual claims, and mathematics used “to bulldoze rather than to clarify”.17
- In my own field, Armstrong, Moffat, Webber and Zobel took a decade of ad-hoc retrieval results reported at SIGIR and CIKM (1998–2008), lined them up against each other rather than each against its own baseline, and found “little evidence of improvement in ad-hoc retrieval technology over the past decade” — because the baselines were weak and nobody was comparing across papers.18 Dozens of individually reported improvements; no aggregate movement.
- A decade later Ferrari Dacrema, Cremonesi and Jannach did the equivalent for neural recommendation. Of 18 algorithms from top-tier conferences, only 7 could be reproduced with reasonable effort; 6 of those 7 were often beaten by simple nearest-neighbour or graph heuristics, and the last did not consistently beat a well-tuned linear method.19
None of that says mathematics is decoration. This is where I have to be fair to the field and to my own younger self. My ECIR paper needed a collapsed Gibbs sampler, and deriving it was not a costume — without it the model was untrainable at any useful scale. Sophistication is frequently load-bearing. The problem is not that our venues reward rigour; it is that rigour is legible and novelty is not, so under time pressure a reviewer can check the first and can only have an opinion about the second.
What I did about it was not heroic. I adapted. I learned the machinery — Dirichlet processes, Chinese restaurant franchises, samplers, the lot — and I used it to carry the ideas I actually cared about. That is the honest version. I did not fight the system; I learned to speak it well enough that it would let my ideas through. I would give the same advice to a student today, with the same lack of pride about it.
Moscow, March 2013
The idea I carried through was word order.
Topic models at the time — LDA and nearly everything built on it — assumed exchangeability: a document is a bag of words, and the order they arrived in carries no information. This is an assumption of convenience. It makes the mathematics tractable. It is also, obviously, false about language, and my thesis was an extended argument that you can drop it and still compute.7
ECIR 2013 was the 35th European Conference on Information Retrieval, held 24–27 March 2013 in Moscow — the easternmost ECIR ever, organised by Yandex and the Higher School of Economics, in the buildings of the former Red October chocolate factory on an island in the Moskva. It took 287 submissions; the main research track accepted 29%. That March was the coldest Moscow had recorded in 33 years.8 I was there on a student accommodation grant from Yandex, presenting an n-gram topic model for time-stamped documents — a model that tracks how topics move through time without throwing away the order of the words inside them.9
My fifth slide is still online. It gives three reasons the bag-of-words assumption is a problem, and the first one reads:
Same words. Same counts. Opposite events. Every model in the room, mine included, was built on a representation that could not tell them apart, and mine was a partial fix rather than a solution.
What I said out loud, and what is not on any slide, was that word order was probably the future of text mining and NLP. The deck’s written “future work” is far more modest — explore nonparametric methods for n-gram topics over time — and I want that on the record before I take any credit. The grand claim was a spoken remark by a PhD student at the end of a twenty-minute talk.
Marc Najork was at that conference. He was then a Principal Researcher at Microsoft Research, one of the people who had built AltaVista, and he gave one of seven Industry Day keynotes, on social signals in Bing — a talk that ended, per the conference report, by trying “to point out the limitations of social search and dispel its myths”.8 Yandex filmed an interview with him in Moscow that week and put it up on 4 April 2013.10 My memory is that he was in the hall for my talk. I cannot prove that part and I am not going to pretend otherwise; what the record establishes is that he was at that conference, and that I was making a large claim in a room that contained people who had built the systems I had grown up reading about.
So was I right?
Partly. And the way in which I was wrong is more interesting than the way in which I was right.
The direction was right. Four years later, the Transformer arrived, and section 3.5 of Attention Is All You Need is a direct statement of the problem my slide was about:
Self-attention, left alone, is a bag-of-words machine. Order has to be put back in deliberately — first with sine and cosine encodings, later with rotary position embeddings, which are what Llama-style models actually use.1112 Position encoding is not a footnote in modern LLMs; it is a live research area, and how far a model can extrapolate beyond its training length turns out to depend on which scheme you chose.13 The thing every one of these systems has in common, underneath the scale, is that they model the order of the tokens.
The mechanism was wrong. I thought word order would enter models the way it entered mine: as explicit discrete structure — n-grams, phrases, collocations, segment boundaries, with a latent variable deciding where a phrase begins.14 That is not what happened. Order became a continuous geometric property of the representation — a rotation in embedding space — and the explicit phrase machinery I spent years on was simply bypassed. Nobody uses my model. The bet on word order paid off; the way I placed it did not.
And the evidence is messier than my version of the story. In 2021 Sinha and colleagues pre-trained masked language models on sentences with the words randomly shuffled, and found they still did well after fine-tuning on many downstream tasks — including tasks specifically designed to punish models that ignore order. Their conclusion was that “purely distributional information largely explains the success of pre-training”.15 Read narrowly, that is a paper saying word order mattered less than people like me assumed. I think the honest reading is that it indicts our benchmarks as much as our models — which is what the authors themselves say, calling for evaluation sets that require deeper linguistic knowledge. But I am not going to quietly leave it out because it is inconvenient.
Set against it: in 2024 Chen, Chi, Wang and Zhou showed that permuting the order of the premises in a logical reasoning problem — a change that cannot alter the answer — causes accuracy drops of over 30% in state-of-the-art LLMs, and that presenting premises in the order the proof needs them dramatically improves performance.16 That is not a model using word order well. It is a model whose behaviour is hostage to order it has not properly understood. Which is, in a sense, the 2013 slide again, thirteen years and several orders of magnitude of compute later, wearing better clothes.
Being early is worth almost nothing
Three things I have taken from this, none of them flattering.
Being right early is not a credential. It is close to worthless, professionally, and people who trade on it are usually rounding. I did not predict the Transformer. I did not predict attention, or scale, or any of the things that actually mattered. I argued that a convenient assumption was false, which was neither original to me — Wallach had said it in 2006, and better14 — nor sufficient. What the record shows is a student relaying what his own experiments demonstrated, in a room, with more confidence than the results strictly licensed. I would rather be described that way accurately than heroically.
The disillusionment was correct and the conclusion I drew from it was wrong. I was right that the system rewards legible sophistication over illegible novelty; the evidence above is not ambiguous. Where I went wrong, for about a year, was concluding that the ideas therefore did not matter. They do. The mathematics is the visa, not the destination. The students I supervise now get told both halves, in that order, because being told only the first half makes cynics and being told only the second makes casualties.
The assumptions of convenience are where the future is. Every field carries a simplification adopted because it made the equations work, which everyone knows is false and nobody has time to remove. Exchangeability was ours for twenty years. The useful question to ask of any field, including this one today, is not “what is the state of the art” but “what is the thing we all know is wrong and tolerate anyway”. In 2026, in my own area, my candidates are that a document has one meaning, that a benchmark score is a measurement, and that the tokenizer is a neutral piece of plumbing.20 I could be wrong about all three. Someone should stand up in a room and say so.
Sources
- Brian Pinkerton, Finding What People Want: Experiences with the WebCrawler, Proceedings of the Second International World Wide Web Conference, Chicago, 17–20 October 1994. Author’s own copy: thinkpink.com (plain HTTP; the site has been online since the 1990s and the page was last modified in 2001).
- Brian Pinkerton, WebCrawler: Finding What People Want, doctoral dissertation, University of Washington, Department of Computer Science & Engineering, 2000; supervisory committee co-chairs Edward Lazowska and John Zahorjan. The quotation is from the Preface, p. v; the abstract describes WebCrawler as “the first comprehensive full-text search engine for the World-Wide Web”. thinkpink.com/bp/Thesis/Thesis.pdf
- Lawrence Page, Sergey Brin, Rajeev Motwani and Terry Winograd, The PageRank Citation Ranking: Bringing Order to the Web, Technical Report 1999-66, Stanford InfoLab, 11 November 1999. The Stanford InfoLab publication server (ilpubs.stanford.edu:8090/422/) was returning a database error when this entry was checked on 12 August 2026; the bibliographic details above are the standard ones and are given so the report can be located elsewhere.
- Sergey Brin and Lawrence Page, The Anatomy of a Large-Scale Hypertextual Web Search Engine, Proceedings of the Seventh International World Wide Web Conference (WWW7), Brisbane, April 1998. Source of the “very little academic research has been done” remark, in the abstract. infolab.stanford.edu
- Primary documents for the 2007–2009 Tata Steel R&D placement — the signed letter of appreciation listing the multithreaded web crawler among five projects, the 89-page B.Tech thesis, and Animesh Bisoee’s Hindustan Times report on the localised search engine (Jamshedpur, August 2008) — are reproduced in full on the undergraduate record.
- Shoaib Jameel, Wai Lam, Ching-man Au Yeung and Sheaujiun Chyan, An Unsupervised Ranking Method Based on a Technical Difficulty Terrain, CIKM 2011, pp. 1989–1992, Glasgow. Winner, Best Paper Award, Beijing–Hong Kong International Doctoral Forum 2011. doi.org/10.1145/2063576.2063872 · the Mario story in full.
- Mohammad Shoaib Jameel, Latent Probabilistic Topic Discovery for Text Documents Incorporating Segment Structure and Word Order, PhD thesis, Systems Engineering and Engineering Management, The Chinese University of Hong Kong, July 2014. Full thesis (PDF, 354pp). The related SIGIR paper is Shoaib Jameel and Wai Lam, An Unsupervised Topic Segmentation Model Incorporating Word Order, SIGIR 2013, pp. 203–212, Dublin (doi); the journal version is Supervised topic models with word order structure for document classification and retrieval learning, Information Retrieval Journal 18(4), 2015, pp. 283–330 (doi).
- Pavel Serdyukov, Pavel Braslavski and Jaap Kamps, Conference Report: ECIR 2013 — 35th European Conference on Information Retrieval, ACM SIGIR Forum 47(2), December 2013, pp. 41–57. Source of the dates (24–27 March 2013), the Yandex/HSE organisation, the Digital October venue in the former Red October chocolate factory, the 287 submissions and 29% main-track acceptance rate, the coldest March in 33 years, and the Industry Day programme — chaired by Yandex CTO Ilya Segalovich, held Wednesday 27 March in parallel with the technical tracks, at which “Marc Najork form [sic] Microsoft Research” spoke on social search. sigir.org
- Shoaib Jameel and Wai Lam, An N-Gram Topic Model for Time-Stamped Documents, ECIR 2013, LNCS 7814, pp. 292–304, Moscow (doi). The quoted line is slide 5 of the conference presentation itself, archived at Semantic Scholar; slide 37 gives the deck’s stated future work. Attendance was supported by an ECIR-2013 student accommodation grant from Yandex.
- Yandex, «Марк Найорк о прошлом и будущем поиска» (“Marc Najork on the past and future of search”), published 4 April 2013. The description states that Najork is a Principal Researcher at Microsoft Research, was among those who worked on AltaVista, and took part in Industry Day at ECIR-2013. youtube.com
- Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser and Illia Polosukhin, Attention Is All You Need, NeurIPS 2017; arXiv:1706.03762, first submitted 12 June 2017. The quotation is the opening sentence of §3.5, “Positional Encoding”. arxiv.org
- Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen and Yunfeng Liu, RoFormer: Enhanced Transformer with Rotary Position Embedding, arXiv:2104.09864, 2021. Rotary position embedding encodes absolute position with a rotation matrix such that the attention inner product depends only on relative position; it is the scheme used by Llama-style models. arxiv.org
- Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das and Siva Reddy, The Impact of Positional Encoding on Length Generalization in Transformers, NeurIPS 2023; arXiv:2305.19466. Compares absolute, T5-relative, ALiBi, rotary and no-positional-encoding decoder-only Transformers on length generalization. arxiv.org
- Hanna M. Wallach, Topic Modeling: Beyond Bag-of-Words, ICML 2006, pp. 977–984. The two other direct antecedents of my ECIR model, both cited in the Moscow deck, are Xuerui Wang and Andrew McCallum, Topics over Time: A Non-Markov Continuous-Time Model of Topical Trends, KDD 2006, pp. 424–433 — the model mine was benchmarked against — and Xuerui Wang, Andrew McCallum and Xing Wei, Topical N-grams: Phrase and Topic Discovery, with an Application to Information Retrieval, ICDM 2007, pp. 697–702.
- Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams and Douwe Kiela, Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for Little, EMNLP 2021, pp. 2888–2913. Note the scope: this concerns masked-language-model pre-training followed by fine-tuning on downstream NLU tasks, not autoregressive generation. aclanthology.org
- Xinyun Chen, Ryan A. Chi, Xuezhi Wang and Denny Zhou, Premise Order Matters in Reasoning with Large Language Models, ICML 2024; arXiv:2402.08939. Permuting premises — which cannot change the correct answer — causes accuracy drops of over 30%; the paper also releases R-GSM, a reordered variant of GSM8K. arxiv.org
- Zachary C. Lipton and Jacob Steinhardt, Troubling Trends in Machine Learning Scholarship, arXiv:1807.03341, July 2018; later published in ACM Queue 17(1), 2019. “Mathiness” is their third named pattern, defined in the abstract as “the use of mathematics that obfuscates or impresses rather than clarifies”; the term is borrowed from Paul Romer. arxiv.org
- Timothy G. Armstrong, Alistair Moffat, William Webber and Justin Zobel, Improvements That Don’t Add Up: Ad-Hoc Retrieval Results Since 1998, CIKM 2009, pp. 601–610, Hong Kong. Analyses TREC Ad-Hoc, Web, Terabyte and Robust results reported at SIGIR (1998–2008) and CIKM (2004–2008). doi.org/10.1145/1645953.1646031
- Maurizio Ferrari Dacrema, Paolo Cremonesi and Dietmar Jannach, Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches, RecSys 2019 (Best Long Paper); arXiv:1907.06902. 18 algorithms examined, 7 reproducible with reasonable effort, 6 of those 7 often outperformed by simple heuristics. An extended analysis appeared as A Troubling Analysis of Reproducibility and Progress in Recommender Systems Research, ACM TOIS, 2021. arxiv.org
- On the tokenizer specifically, see MyTh 002 and the accompanying interactive measurement: the same sentence of the Universal Declaration of Human Rights costs 33 tokens in English and 305 in Amharic.
Checked against primary sources on 12 August 2026. Three things are worth flagging rather than burying. First, my recollection that Marc Najork was in the hall for my talk is a recollection and nothing more — the conference report8 places him at ECIR 2013 as an Industry Day speaker, and Industry Day ran on the final morning in parallel with the technical tracks, but no record establishes who sat in which session. Second, the claim that word order was the future of text mining was said out loud and is not in the slide deck; the deck’s written future work is narrower, and I have quoted it above so you can see the difference.9 Third, the strongest published evidence against my position — Sinha et al.15 — is included deliberately, and I have tried to state it at its strongest rather than at its most convenient. The Stanford InfoLab server hosting the PageRank technical report was down when this was written, which is why that citation carries bibliographic detail instead of a working link. Corrections to M.S.Jameel@southampton.ac.uk.