How does an inverted index make full-text search fast?
basicAn inverted index maps each term to the list of documents containing it (plus positions and frequencies), so a search looks up terms directly instead of scanning every document. Each Elasticsearch shard is a Lucene index made of immutable segments, each with its own inverted index.
- Indexing: text is analyzed into terms, then added to the term dictionary and postings lists.
- Query: look up each query term, intersect or union postings lists, then score.
- Segments are immutable: updates write a new document version and mark the old one deleted; background merges purge deleted docs.
- Doc values (columnar, on disk) serve sorting and aggregations, because the inverted index is poor at "value of field for doc".
- Is an update in place? No, it is delete plus reindex of the whole document.
- Why can't you sort on an analyzed
textfield? It has no doc values; use akeywordsub-field.