Escapes From the Tree
A tree gives each thing one path, so it can collocate one facet and must scatter the rest — the citation order problem. Five mechanisms let a reader reach a thing by a facet other than the first-cited one. Each keeps a single physical location and adds a second route to it, and each pays for the route in a different currency: a polyhierarchy gives a node more than one parent; a link or alias is a second name that resolves to the first; postcoordination defers the combination of facets to query time; rotation generates every citation order in advance; and a lattice stores every combination of facets as a node of its own. What the evidence says people actually do comes last, because it is not what the mechanisms predict.
Polyhierarchy#
ISO 25964 defines a monohierarchical structure as one “in which each concept can have only one broader concept at the level immediately above”, with the example that “the concept of pianos cannot be listed under keyboard instruments as well as under stringed instruments; a choice has to be made of one of these concepts to determine its placing.” A polyhierarchical structure is one “in which each concept can have more than one broader concept”. ANSI/NISO Z39.19 (the standard of the National Information Standards Organization (NISO)) allows it on the same terms — “Some concepts belong, on logical grounds, to more than one category” — and the World Wide Web Consortium’s Simple Knowledge Organization System (SKOS) builds it in: “a SKOS concept can be attached to several broader concepts at the same time. For example, a concept ex:dog could have both ex:mammals and ex:domesticatedAnimals as broader concepts.” Structurally the tree becomes a directed acyclic graph: the single-parent rule is dropped and nothing else changes. Christopher Alexander gave the general form in 1965. “A collection of sets forms a tree if and only if, for any two sets that belong to the collection either one is wholly contained in the other, or else they are wholly disjoint”, whereas it forms a semilattice “if and only if, when two overlapping sets belong to the collection, the set of elements common to both also belongs to the collection.” The difference in capacity is the point of his essay: “A tree based on 20 elements can contain at most 19 further subsets of the 20, while a semilattice based on the same 20 elements can contain more than 1,000,000 different subsets.”
This is how Wikipedia answers the vendor/kind question — “A category can have more than one parent category”, and Category:Microsoft software sits under both Microsoft and Software by company — and how the Nielsen Norman Group recommends a shop answer it: a Nintendo Switch listed under both Video Games and Electronics, because “different users assume that either Video Games or Electronics are the natural ‘home’” for it. The cost is that the tree stops being navigable if the second parent is used freely. The same article: “every single place where a particular item could sit would swell each menu with tons of items, and add significant cognitive strain for users”, and a breadcrumb has to pick one “canonical path” that may not be the path the reader took. Hedden’s working limits are that a polyhierarchy “usually involves only two broader concepts, not more”, that it is legitimate only “when the concept’s relationship is correctly and inherently hierarchical in both of its cases” (see Logic of Division), and that “More than 2-3 polyhierarchies across an entire faceted taxonomy should be a cause for review”. Wikipedia’s own guideline against narrow intersection categories is the same limit from the other side: “if an article is in category ‘A’ and in category ‘B’ – a category A and B does not necessarily need to be created for this article.”
Links, aliases and canonical paths#
Unix had the second name from the start. Ritchie and Thompson, 1974: “The same nondirectory file may appear in several directories under possibly different names. This feature is called linking … a file does not exist within a particular directory; the directory entry for a file consists merely of its name and a pointer to the information actually describing the file.” Files could have many paths; directories could not, for the reason quoted on the section page: a rooted tree is what makes it possible to tell when a directory has been cut off from the root. Hard links “do not allow references across physical file systems”, so the Berkeley Software Distribution (BSD) of 1982–83 added symbolic links “similar to the scheme used by Multics”, each “implemented as a file that contains a pathname” — which is why a symbolic link can cross file systems and point at a directory, and why it can dangle when the target moves. The Presto paper from Xerox’s Palo Alto Research Center gave the verdict of everyday use in 1999: “The use of links and aliases is testament to this problem, but at the same time is sufficiently complicated that it often extends no further than desktop shortcuts.”
On the web the second name is a redirect, and the discipline around it is canonicalization: “the process of selecting the representative –canonical– URL of a piece of content”, in Google’s definition, needed because among other things “the results of sorting and filtering functions of a category page” produce many URLs for one page. A redirect is “a strong signal that the target of the redirect should become canonical”, rel="canonical" is the same signal declared in the page, and either way exactly one path is the page’s address and the rest are pointers. Hugo’s aliases front matter is this mechanism for a static site — a small HTML file per old URL carrying <meta http-equiv="refresh"> — and it is the repair for a rename described in Content Organization. MediaWiki’s #REDIRECT [[pagename]] is the same idea for a wiki, “useful if a particular article is referred to by multiple names”. Shirky’s parable of Yahoo shows what happens when a directory adds the second name reluctantly: a category marked with an @ “isn’t ‘really’ in the category Entertainment”, and “A URL can only appear in three places. That’s the Yahoo rule.”
A link costs nothing to create and something to maintain: it is second-class by design (the target is canonical, the alias is not), it breaks silently when the target moves, and — Presto’s point — a scheme that needs many of them is a scheme whose citation order is wrong for its readers.
Postcoordination#
The postcoordinate answer is to stop combining facets at filing time at all. Mortimer Taube defined coordinate indexing in 1951 as “the analysis of any field of information into a set of terms and the combination of these terms in any order to achieve any desired degree of detail in either indexing or selection”, and Svenonius’s history explains what was at stake: “whether subject languages need to employ a syntax at all. This issue was first raised in 1951 when the potential of the computer in information retrieval was first discussed. Mortimer Taube, an early visionary, proposed replacing the relatively complex grammatical syntax used to construct subject headings with a simple logical syntax”, Boolean AND, OR and NOT. Under that syntax microsoft AND spreadsheet has no first term.
Everything that indexes by several independent labels is a descendant. Faceted navigation, in Hearst’s definition, is a collection where “multiple labels are assigned to each item, as opposed to a strictly hierarchical system in which items are placed into single categories or folders”, and it “generally works best if the facets are conceptually orthogonal”. Her group’s 2003 study put 32 art-history students in front of 35,000 images with facets and found that “90% of the participants preferred the metadata approach overall”. Endeca, founded in 1999, sold the same thing to retailers as Guided Navigation, a way “to browse information, or to refine long lists of search results, along multiple dimensions, aka facets”: a reader composes (Time: 18th Century) + (Country: France) + (Topical Subject: History) in whatever order they think of it. The Online Computer Library Center’s faceted derivative of the Library of Congress headings, Faceted Application of Subject Terminology (FAST), splits each precoordinated heading into eight facets so that “the user can mix and match”. Tags are the uncontrolled end of the same line — Vander Wal’s folksonomy (2004) is “the result of personal free tagging of information and objects (anything with a URL) for one’s own retrieval” — and Shirky’s list of when a fixed scheme works is really a list of when to prefer this end: “Small corpus, Formal categories, Stable entities, Restricted entities, Clear edges” and “Expert catalogers, Authoritative source of judgment, Coordinated users, Expert users”; invert every item and “ontology is going to be a bad strategy.”
The costs are on record too, from the institution with the most precoordinated strings in the world. The Library of Congress’s 2007 review concluded that “pre-coordinated strings provide context, which is needed for ‘disambiguation, suggestibility, and precision’ and browsability”, and that “Post-coordinated terms have serious limitations for recall, precision, understanding, and relevance ranking.” Svenonius, fifty years after Taube, called the question “still unresolved”, noting “a common criticism leveled against postcoordinate languages is their inability to express relationships more specific than those expressible by” the Boolean operators. Rafferty’s survey of tagging lists the price of uncontrolled labels — no “synonym and homonym control, precision and hierarchy” — and Samuel’s 2019 retrospective records that tagging lost not to hierarchy but to the feed: “The rise of Twitter and Facebook—and especially, the arrival of the Facebook home feed—gave online news consumers an easier option.”
Operating systems have tried to build postcoordination in below the file system. Gifford’s Semantic File Systems (1991) added “virtual directories” whose “names are interpreted as queries”, so /sfs/owner:/smith/text:/resume and /sfs/field:/text:/semantic/owner:/jones are both paths, with the attributes in whichever order the reader thinks of them. BeOS shipped indexed attributes and live queries in its file system. Microsoft’s WinFS (shown in 2003) was to let “an item … reside in multiple folders without duplicating the actual data” and was withdrawn in June 2006: “we are not pursuing a separate delivery of WinFS, including the previously planned Beta 2 release.” Seltzer and Murphy’s 2009 paper Hierarchical File Systems are Dead proposed “a tagged, search-based” namespace instead. None displaced the tree.
Rotation#
Rotation keeps precoordination but generates every citation order, so that a reader arriving by any facet finds an entry. Ranganathan’s chain procedure (1938) derives an index entry from each link of a class number, and PRECIS — the Preserved Context Index System that Derek Austin built for the British National Bibliography from 1968 — took one string of terms with role operators and, by a process known as shunting, produced an entry led by each term with the rest preserved as context: India / Libraries / Computerisation yields entries under India, Libraries and Computerisation. Voit’s TagTrees (2011) do this for a file system with symbolic links: an item tagged Bob and MyProject “may be found along any one of four possible paths Bob, MyProject, Bob/MyProject, and MyProject/Bob”, because “The TagTree folder hierarchy consists of one folder path for each permutation of the tags associated with the item”, so that “it is not necessary to remember a single strict series of folder names.”
The cost is the factorial. Gray’s data cube is rotation for aggregates and states the arithmetic: CUBE “requires generating the power set (set of all subsets) of the aggregation columns”, giving “2^N-1 super-aggregate values” for N dimensions, while ROLLUP — one fixed order — is what you use when the facets are genuinely nested, since “a cube on these three attributes would be meaningless” for year, week and day. Formal concept analysis (Wille, 1982) is the full version: every combination of attributes that some set of objects shares becomes a node, and “The information contained in the formal context is preserved” in a concept lattice. It also shows why nobody browses one: the lattice contains every citation order because it contains every intersection. PRECIS was retired from the bibliography in 1996; Šauperl’s review notes that precoordinated systems “require investments in highly skilled intellectual work, and are therefore expensive and difficult to maintain.”
Luhmann’s slip box is rotation refused. Each slip got a number “which is easily seen … and that we never change”, branching by insertion (“57/12 can then be continued with 57/13 … it can be supplemented … by 57/12a or 57/12b”); he chose to “decide against the systematic ordering in accordance with topics and sub-topics and choose instead a firm fixed place”, and reached the other facets by links and “a register of keywords that we constantly update”. That is one canonical path plus a second mechanism, which is where the evidence points.
What the evidence says#
Studies of people organising their own files keep finding the tree in use, with search as the fallback. Bergman and colleagues asked 296 people to retrieve 1,131 of their files and logged 5,035 navigation steps: “Folder navigation is the main way that personal computer users retrieve their own files”, structures were shallow (mean depth 2.86), and navigation found 94% of files in 14.76 seconds on average. Their literature review is blunter: “regardless of search engine quality, there was a strong preference for navigation. Search was predominantly used as a last resort only when users could not remember the location of a file.” Their 2013 comparison gave 75 Gmail users labels-as-folders and labels-as-tags and found “a strong preference for folders over tags for both storage and retrieval”; “when multiple classification was used for storage, it was only marginally used for retrieval”. Jones offered fourteen people a hypothetical perfect search in exchange for their folders, and “13 of 14 participants gave a resounding ‘No!’”, giving reasons that are about the tree as a plan, not a retrieval device: “Folders help me see the relationship between things”, “Folders remind me what needs to be done”. The one large study on the other side is email: Whittaker’s team logged 345 users and 85,000 refinding actions and found that “People who create complex folders indeed rely on these for retrieval, but these preparatory behaviors are inefficient and do not improve retrieval success. In contrast, both search and threading promote more effective finding.” Practitioners read the same split. Hedden’s knowledge graphs still “include a taxonomy, thesaurus, or set of controlled vocabularies to provide consistent labeling”, and Arango, using a language model to cluster his own writing, kept the last step: “The AI suggested possible groupings, but I defined the final themes.”
Put together: postcoordination wins for retrieval by many untrained readers over large collections, which is Shirky’s list and the shop’s product catalogue; a single fixed hierarchy wins for a person’s own material and for anything that has to be a plan or a taxonomy; and every system that has to serve both keeps one canonical path and adds exactly one of the mechanisms above for the other facet.
Check#
For a tree that has to serve readers arriving by more than one facet:
- Choose the first-cited facet by the is-a test if the tree is a taxonomy, and by the readers’ purpose — a card sort — if it is not.
- Apply one basis of division per level, to every sibling.
- Give the second facet a route that is not a second copy: an index page, a tag, a redirect, or a second parent, and no more than two parents.
- When a thing moves, leave an alias at the old path, and run whatever checks internal links; Hugo does not.
Sources#
- ISO 25964-1:2011, definitions 2.34 and 2.42 (preview); ANSI/NISO Z39.19-2005, §8.3.4 (PDF); SKOS Primer, §2.3.1
- Alexander, C., “A City is Not a Tree”, Architectural Forum 122(1–2), 1965 (text)
- Laubheimer, P., “Polyhierarchies Improve Findability for Ambiguous IA Categories” (IA: information architecture), Nielsen Norman Group, 2018; Hedden, H., “Polyhierarchy in Taxonomies”, 2022; Wikipedia, Overcategorization
- Ritchie, D. M. and Thompson, K., “The UNIX Time-Sharing System”, Communications of the ACM 17(7), 1974 (PDF); McKusick, Joy, Leffler and Fabry, “A Fast File System for UNIX”, §5.3 (PDF)
- Dourish, P., Edwards, W. K., LaMarca, A. and Salisbury, M., “Presto: An Experimental Architecture for Fluid Interactive Document Spaces”, ACM Transactions on Computer-Human Interaction 6(2), 1999 (PDF)
- Google Search Central, What is URL canonicalization and How to specify a canonical URL; Hugo, URL management; MediaWiki, Help:Redirects
- Shirky, C., “Ontology is Overrated: Categories, Links, and Tags”, 2005
- Taube, M., “Coordinate Indexing of Scientific Fields”, paper to the American Chemical Society, 4 September 1951, as quoted at History of Information; Svenonius, E., The Intellectual Foundation of Information Organization (MIT Press, 2000), pp. 188–192, excerpted in the Library of Congress report Pre- vs. Post-Coordination and Related Issues, 2007
- Hearst, M. A., “UIs for Faceted Navigation”, 2008; Yee, Swearingen, Li and Hearst, “Faceted Metadata for Image Search and Browsing”, 2003; Papa, S., “A Primer on Faceted Navigation and Guided Navigation”, 2004; Wikipedia, Faceted Application of Subject Terminology
- Vander Wal, T., “Folksonomy”, 2007; Rafferty, P. M., “Tagging”, ISKO Encyclopedia of Knowledge Organization (the ISKO is the International Society for Knowledge Organization); Samuel, A., “What Happened to Tagging?”, JSTOR Daily, 2019
- Gifford, D. K., Jouvelot, P., Sheldon, M. A. and O’Toole, J. W., “Semantic File Systems”, ACM Operating Systems Review, October 1991 (PDF); Giampaolo, D., Practical File System Design with the Be File System (Morgan Kaufmann, 1999) (PDF); Clark, Q., “WinFS Update”, 23 June 2006; Seltzer, M. and Murphy, N., “Hierarchical File Systems are Dead”, 2009
- Austin, D., PRECIS: A Manual of Concept Analysis and Subject Indexing (Council of the British National Bibliography, 1974); overview at Librarianship Studies; Voit, K., Andrews, K. and Slany, W., “TagTree: Storing and Re-finding Files Using Tags”, 2011
- Gray, J. et al., “Data Cube: A Relational Aggregation Operator Generalizing Group-By, Cross-Tab, and Sub-Totals”, Data Mining and Knowledge Discovery 1(1), 1997 (PDF); Ganter, B., “Formal Concept Analysis”, 2008; Šauperl, A., “Precoordination or not?”, Journal of Documentation 65(5), 2009 (abstract)
- Luhmann, N., “Communicating with Slip Boxes” (1981), trans. M. Kuehn
- Bergman, O., Whittaker, S., Sanderson, M., Nachmias, R. and Ramamoorthy, A., “The effect of folder structure on personal file navigation”, Journal of the American Society for Information Science and Technology 61(12), 2010 (PDF); Bergman, Gradovitch, Bar-Ilan and Beyth-Marom, “Folder versus tag preference in personal information management”, same journal, 64(10), 2013 (abstract); Jones, W., Bruce, H., Foxley, A. and Munat, C. F., “Planning Personal Projects and Organizing Personal Information”, 2006; Whittaker, S., Matthews, T., Cerruti, J., Badenes, H. and Tang, J., “Am I wasting my time organizing email?”, CHI 2011 (ACM)
- Hedden, H., “Knowledge Graphs and Taxonomies”, 2023; Arango, J., “Using AI as an Assistant to Organize Content”, 2024