Torus (Torus v0.7.0)
View SourceTorus is a plug-and-play Elixir library that seamlessly integrates PostgreSQL's search into Ecto, streamlining the construction of advanced search queries. See live demo for examples.
Usage
The package can be installed by adding torus to your list of dependencies in mix.exs:
def deps do
[
{:torus, "~> 0.7"}
]
endThen, in any query, you can (for example) add a prefixed full-text search:
import Torus
# ...
Post
# ... your complex query
|> Torus.full_text([p], [p.title, p.body], "uncove hogwar")
|> select([p], p.title)
|> Repo.all()
["Uncovered hogwarts"]See full_text/5 for more details.
7 types of search:
Pattern matching: Searches for a specific pattern in a string.
iex> insert_posts!(["Wand", "Magic wand", "Owl"]) ...> Post ...> |> Torus.similar_to([p], [p.title], "(Wan|Ow)%") ...> |> select([p], p.title) ...> |> Repo.all() ["Wand", "Owl"]Use it for fast prefix-search when semantics of the data you search through live in its characters. For example phone number, invoice number, email, filename, etc.
See
like/5,ilike/5, andsimilar_to/5for more details.Similarity: Searches for records that closely match the input text using trigram distance.
insert_posts!(["Hogwarts Secrets", "Quidditch Fever", "Hogwart’s Secret"]) Post |> Torus.similarity([p], [p.title], "hoggwarrds") |> limit(2) |> select([p], p.title) |> Repo.all() ["Hogwarts Secrets", "Hogwart’s Secret"]Use it for fuzzy matching and catching typos in short text fields, such as names or titles. Works best with short strings.
See
similarity/5for more details.Full text: Uses term-document matrix vectors, enabling efficient querying and ranking based on term frequency. Supports prefix search and is great for large datasets to quickly return relevant results. See PostgreSQL Full Text Search for internal implementation details.
insert_post!(title: "Hogwarts Shocker", body: "A spell disrupts the Quidditch Cup.") insert_post!(title: "Diagon Bombshell", body: "Secrets uncovered in the heart of Hogwarts.") insert_post!(title: "Completely unrelated", body: "No magic here!") Post |> Torus.full_text([p], [p.title, p.body], "uncov hogwar") |> select([p], p.title) |> Repo.all() ["Diagon Bombshell"]Use it when you don't care about spelling, the documents are long, you need multi-column search with weights, or if you need to order the results by rank.
See
full_text/5for more details.BM25 full text: Modern BM25 ranking algorithm for superior relevance scoring using the pg_textsearch extension. BM25 generally provides better ranking than traditional built-in TF-IDF full text search and is optimized for top-k queries.
insert_post!(title: "Potion Class Notes", body: "Wiggenweld potion heals wounds.") insert_post!(title: "Complete Potion Encyclopedia", body: "Edurus potion grants protection. Focus potion improves concentration. Maxima potion amplifies spells. Thunderbrew potion creates explosions.") insert_post!(title: "Combat Guide", body: "Use Wiggenweld potion to heal during goblin fights.") Post |> Torus.bm25([p], p.body, "wiggenweld potion") |> limit(2) |> select([p], p.title) |> Repo.all() ["Potion Class Notes", "Combat Guide"]Use it when you need state-of-the-art relevance ranking for single-column search, especially with LIMIT clauses. Requires PostgreSQL 17+.
See
bm25/5and the BM25 Search Guide for detailed setup instructions and examples.Semantic Search: Understands the contextual meaning of queries to match and retrieve related content utilizing natural language processing. Read more about semantic search in Semantic search with Torus guide.
insert_post!(title: "Hogwarts Shocker", body: "A spell disrupts the Quidditch Cup.") insert_post!(title: "Diagon Bombshell", body: "Secrets uncovered in the heart of Hogwarts.") insert_post!(title: "Completely unrelated", body: "No magic here!") embedding_vector = Torus.to_vector("A magic school in the UK") Post |> Torus.semantic([p], p.embedding, embedding_vector) |> select([p], p.title) |> Repo.all() ["Diagon Bombshell"]Use it when you need to understand intent and handle synonyms.
See
semantic/5for more details.Hybrid Search: Combines multiple search techniques (e.g., keyword and semantic) to leverage their strengths for more accurate results. Fusion happens in a single SQL query using Reciprocal Rank Fusion: each branch ranks its best rows, and rows are merged by summing
weight * 1.0 / (k + rank)across branches.insert_post!(title: "Hogwarts Shocker", body: "A spell disrupts the Quidditch Cup.") insert_post!(title: "Diagon Bombshell", body: "Secrets uncovered in the heart of Hogwarts.") insert_post!(title: "Completely unrelated", body: "No magic here!") Post |> Torus.hybrid([p], [ full_text: {[p.title, p.body], "uncov hogwar"}, similarity: {[p.title], "hogwarts"} ]) |> select([p], p.title) |> Repo.all() ["Diagon Bombshell", "Hogwarts Shocker", "Completely unrelated"]Use it when no single search type is good enough - typically combining keyword search (
full_textorbm25) withsemanticsearch for RAG and retrieval pipelines.See
hybrid/4and the Hybrid search guide for more details.3rd Party Engines/Providers: Utilizes external services or software specifically designed for optimized and scalable search capabilities, such as Elasticsearch or Algolia.
You can see all of the above search types in action on the live demo page.
Highlighting matches
Searches can highlight their matches in the results - pass highlight: [key: column] to full_text, bm25, similarity, ilike, like, or a hybrid branch's options, and the search's own term and options are reused:
Post
|> Torus.full_text([p], [p.title, p.body], "shocker", highlight: [title: p.title])
|> Repo.all()
[%Post{title: "Hogwarts <b>Shocker</b>", ...}]For full control (custom terms, snippets), use highlight/3 directly in select/select_merge:
Post
|> Torus.full_text([p], [p.title, p.body], "shocker")
|> select([p], Torus.highlight(p.title, "shocker"))
|> Repo.all()
["Hogwarts <b>Shocker</b>"]Optimizations and relevance
Torus is designed to be as efficient and relevant as possible from the start. But handling large datasets and complex search queries tends to be tricky. The best way to combine these two to achieve the best result is to:
- Create a query that returns as relevant results as possible (by tweaking the options of search function). If there is any option missing - feel free to open an issue/contribute back with it, or implement it manually using fragments.
- Test its performance on real production data - maybe it's good enough already?
- If it's not:
- See optimization sections for your search type in
Torusdocs - Inspect your query using
Torus.QueryInspector.tap_substituted_sql/3orTorus.QueryInspector.tap_explain_analyze/3 - According to the above SQL - add indexes for the queried rows/vectors
- See optimization sections for your search type in
Debugging your queries
Torus offers a few helpers to debug, explain, and analyze your queries before using them on production. See Torus.QueryInspector for more details.
Torus support
For now, Torus supports pattern match, similarity, full-text (TF-IDF and BM25), semantic, and hybrid search, with plans to expand support further. These docs will be updated with more examples on which search type to choose and how to make them more performant (by adding indexes or using specific functions).
Summary
Full text
BM25 ranked full-text search using the pg_textsearch extension.
Full text search with rank ordering. Accepts a list of columns to search in. A list of columns
can either be a text or tsvector type. If tsvectors are passed make sure to set
stored: true.
Highlights matches of term in qualifier using PostgreSQL
ts_headline.
Hybrid
Hybrid search: fuses several search strategies into a single ranked query using
Reciprocal Rank Fusion
(RRF). Each branch runs as an independent ranked subquery, keeps its limit best rows,
and the results are merged by summing weight * 1.0 / (k + rank) per row across branches.
Pattern matching
Case-insensitive pattern matching search using
PostgreSQL ILIKE operator.
Case-sensitive pattern matching search using PostgreSQL LIKE operator.
Removes all like/ilike special characters from the term, so it can be used in further pattern-match searches.
Similar to like/5, except that it interprets the pattern using the SQL standard's
definition of a regular expression. SQL regular expressions are a curious cross between
LIKE notation and common (POSIX) regular expression notation. See
PostgreSQL SIMILAR TO
The substring function with three parameters provides extraction of a substring that matches an SQL regular expression pattern. The function can be written according to standard SQL syntax
Semantic
Calls the specified embedding module's embedding_model/1 function to retrieve the model name.
Semantic search using pgvector extension to compare vectors. See Semantic search guide for more info.
Same as to_vectors/2, but returns the first vector from the list.
Takes a list of terms (binaries) and embedding module's specific options and passes them to embedding_module generate/2 function.
Similarity
Searches for records that closely match the input text using trigram distance. Ideal for fuzzy matching and catching typos in short text fields.
Full text
BM25 ranked full-text search using the pg_textsearch extension.
BM25 is a modern ranking function that generally provides better relevance than traditional
TF-IDF (used by full_text/5). It's particularly effective for top-k queries with LIMIT clauses
due to Block-Max WAND optimization.
For detailed usage examples, performance tips, and migration guide, see the BM25 Search guide.
Requirements
- Requires the
pg_textsearchextension to be installed - PostgreSQL 17+ only
- Requires a BM25 index on the search column
- Single column only - unlike
full_text/5, BM25 indexes work on one column at a time - Language is set at index creation - use
text_configin the indexWITHclause
defmodule YourApp.Repo.Migrations.CreatePgTextsearchExtension do
use Ecto.Migration
def change do
execute "CREATE EXTENSION IF NOT EXISTS pg_textsearch", "DROP EXTENSION IF EXISTS pg_textsearch"
# Create BM25 index with language configuration
execute """
CREATE INDEX posts_body_bm25_idx ON posts
USING bm25(body) WITH (text_config='english')
""", "DROP INDEX posts_body_bm25_idx"
end
endOptions
:order- Ordering of results. Note that BM25 returns negative scores (lower is better)::asc(default) - orders by score ascending (best matches first):desc- orders by score descending (worst matches first):none- no ordering applied
:index_name- Explicit index name. Required when usingscore_threshold.:score_key- Atom key to select the BM25 score into the result map. The score is merged viaselect_merge/3, so the query needs to select a map before callingbm25/5(e.g.select([p], %{body: p.body})).:none(default) - score is not selectedatom- selects score as this key
:score_threshold- Post-filter results by BM25 score (applied after ORDER BY). Since scores are negative and lower is better, use negative thresholds (e.g.,-3.0keeps only results with score < -3.0, i.e., scores like -4.0, -5.0 which are better matches). May return fewer results than LIMIT.:pre_filter- Whether to exclude non-matching rows.false(default) - no pre-filteringtrue- adds aWHERE score < 0clause to exclude non-matches
:highlight- a keyword list of result keys to columns to highlight the term's matches in, e.g.highlight: [body: p.body]. Seehighlight/3.
Examples
Basic search - returns top 10 most relevant posts:
Post
|> Torus.bm25([p], p.body, "database search")
|> limit(10)
|> select([p], p.body)
|> Repo.all()With score selection (select a map before calling, so the score has somewhere to merge into):
Post
|> select([p], %{body: p.body})
|> Torus.bm25([p], p.body, "database", score_key: :relevance)
|> limit(5)
|> Repo.all()
# => [%{body: "...", relevance: -2.5}, ...]With WHERE clause pre-filtering:
Post
|> where([p], p.category_id == 123)
|> Torus.bm25([p], p.body, "database")
|> limit(10)
|> Repo.all()With score threshold (post-filtering, may return fewer than LIMIT, index_name is required):
Post
|> Torus.bm25([p], p.body, "database", score_threshold: -5.0, index_name: "posts_body_idx")
|> limit(10)
|> Repo.all()When to use bm25/5 vs full_text/5
Use bm25/5 when:
- You need better relevance ranking than TF-IDF
- You need faster search with large datasets
- You have large result sets with LIMIT (top-k queries)
- Single column search is sufficient
- You're on PostgreSQL 17+
Use full_text/5 when:
- You need multi-column search with different weights per column
- You want to use stored tsvector columns
- You're on PostgreSQL < 17
- You need the
concatfilter type
Index options
BM25 indexes support these parameters in the WITH clause:
text_config- PostgreSQL text search configuration (required). This determines the language/stemming rules. Available configs:'english','french','german','simple'(no stemming), etc. RunSELECT cfgname FROM pg_ts_config;to list all.k1- Term frequency saturation (default: 1.2, range: 0.1-10.0)b- Length normalization (default: 0.75, range: 0.0-1.0)
CREATE INDEX custom_idx ON documents
USING bm25(content)
WITH (text_config='english', k1=1.5, b=0.8);Performance tips
- BM25 is most efficient with
ORDER BY + LIMIT(enables Block-Max WAND optimization) - For filtered searches, create a separate B-tree index on the filter column
- Pre-filtering works best when the filter is selective (<10% of rows)
- Post-filtering with
score_thresholdmay return fewer results than LIMIT
Full text search with rank ordering. Accepts a list of columns to search in. A list of columns
can either be a text or tsvector type. If tsvectors are passed make sure to set
stored: true.
Cleans the term, so it can be input directly by the user. The default preset of settings is optimal for most cases.
Full Text Searching (or just text search) provides the capability to identify natural-language documents that satisfy a query, and optionally to sort them by relevance to the query. The most common type of search is to find all documents containing given query terms and return them in order of their similarity to the query. Notions of query and similarity are very flexible and depend on the specific application. The simplest search considers query as a set of words and similarity as the frequency of query words in the document. Read more in PostgreSQL Full Text Search docs.
Options
:prefix_search- whether to apply prefix search.true(default) - the term is treated as a prefixfalse- only counts full-word matches
:stored- whether to use stored tsvector or not.false(default) - columns (or expressions) passed as qualifiers are of typetexttrue- columns (or expressions) passed as qualifiers are tsvectors
:language- language used for the search. Defaults to"english".:term_function- function used to convert the term tots_query. Can be one of::websearch_to_tsquery(default) - converts term to a tsquery, normalizing words according to the specified or default configuration. Quoted word sequences are converted to phrase tests. The word “or” is understood as producing an OR operator, and a dash produces a NOT operator; other punctuation is ignored. This approximates the behavior of some common web search tools.:plainto_tsquery- converts term to a tsquery, normalizing words according to the specified or default configuration. Any punctuation in the string is ignored (it does not determine query operators). The resulting query matches documents containing all non-stopwords in the term.:phraseto_tsquery- converts term to a tsquery, normalizing words according to the specified or default configuration. Any punctuation in the string is ignored (it does not determine query operators). The resulting query matches phrases containing all non-stopwords in the text.
:rank_function- function used to rank the results.:ts_rank_cd(default) - computes a score showing how well the vector matches the query, using a cover density algorithm. See Ranking Search Results for more details.:ts_rank- computes a score showing how well the vector matches the query.
:rank_weights- a list of weights for each column. Defaults to[:A, :B, :C, :D], padded with:Dwhen there are more than four columns. If provided, it should have at least as many weights as there are columns. A single weight can be either a string or an atom. Possible values are::A- 1.0:B- 0.4:C- 0.2:D- 0.1
:rank_normalization- an integer that specifies whether and how a document's length should impact its rank. The integer option controls several behaviors, so it is a bit mask: you can specify one or more behaviors using|(for example,2|4).0- ignores the document length1(default forts_rank) - divides the rank by 1 + the logarithm of the document length2- divides the rank by the document length4(default forts_rank_cd) - divides the rank by the mean harmonic distance between extents (this is implemented only byts_rank_cd)8- divides the rank by the number of unique words in document16- divides the rank by 1 + the logarithm of the number of unique words in document32- divides the rank by itself + 1
:order- describes the ordering of the results. Possible values are:desc(default) - orders the results by similarity rank in descending order.:asc- orders the results by similarity rank in ascending order.:none- doesn't apply ordering at all.
:filter_type- filter type:or(default) - usesORoperator to combine different column matches. Selecting this option means that the search term won't match across columns.:concat- joins the columns into a single tsvector and searches for the term in the concatenated string containing all columns.:none- doesn't apply any filtering and returns all results.
:empty_return- whether to return all results when the search term is empty.true(default) - returns all results when the search term is empty.false- returns an empty list when the search term is empty.
:highlight- a keyword list of result keys to columns to highlight the term's matches in, e.g.highlight: [title: p.title]. Seehighlight/3.:coalesce- when joining multiple columns viafilter_type: :concat, wraps each column in aCOALESCEfunction to handle NULL values. Only applies withfilter_type: :concatand more than one column - it is ignored otherwise.true(default) - addsCOALESCEfalse- doesn't addCOALESCEto the query. Choose this when you can guarantee that all columns are non-null.
Example usage
iex> insert_post!(title: "Hogwarts Shocker", body: "A spell disrupts the Quidditch Cup.")
...> insert_post!(title: "Diagon Bombshell", body: "Secrets uncovered in the heart of Hogwarts.")
...> insert_post!(title: "Completely unrelated", body: "No magic here!")
...> Post
...> |> Torus.full_text([p], [p.title, p.body], "uncov hogwar")
...> |> select([p], p.title)
...> |> Repo.all()
["Diagon Bombshell"]Optimizations
Store precomputed tsvector in a separate column, add a GIN index to it, and use
stored: true.Add a GIN tsvector index on the column(s) you search in. Use
Torus.QueryInspector.tap_sql/2on your query (with all the options passed) to see the exact search string and add an index to it. For example for nullable title, the GIN index could look like:CREATE INDEX index_gin_posts_title ON posts USING GIN (to_tsvector('english', COALESCE(title, '')));
Highlights matches of term in qualifier using PostgreSQL
ts_headline.
Use it in select/select_merge alongside any search macro. Two highlighting
types are supported via the :type option:
:word(default) - word-based, uses the same term parsing asfull_text/5, so it pairs naturally withfull_text/5,bm25/5, andhybrid/4. It also highlights the exact-word matches ofsimilarity/5(though not its fuzzy matches).:substring- highlights every occurrence of the term as a plain substring (implemented withregexp_replace, the term is regex-escaped). Pairs withilike/5andlike/5, whose%term%patterns match inside words where word-based highlighting finds nothing.
semantic/5 matches aren't lexical, so there is nothing to highlight there.
Highlighting from the search macros
Instead of repeating the term, pass highlight: [result_key: column] directly to
full_text/5, bm25/5, similarity/5, ilike/5, like/5, or a hybrid/4
branch's options - the search's own term and options are reused. ilike/5/like/5 highlight as substrings with the
macro's case sensitivity, stripping %/_ wildcards from the term; the rest
highlight word matches. Only leading/trailing wildcards translate cleanly: a
wildcard in the middle of the term (or an escaped \%) is stripped too, so the
remaining substring may no longer occur in the text and nothing gets highlighted -
for such patterns call highlight/3 yourself with the substring you want marked.
The highlighted value is merged via select_merge/3, so
either use a key that exists on the selected struct (its value is replaced with
the highlighted text) or select a map before the search macro.
iex> insert_post!(title: "Hogwarts Shocker")
...> Post
...> |> Torus.full_text([p], [p.title], "shocker", highlight: [title: p.title])
...> |> Repo.all()
...> |> Enum.map(& &1.title)
["Hogwarts <b>Shocker</b>"]
iex> insert_post!(title: "Hogwarts Shocker")
...> Post
...> |> Torus.ilike([p], [p.title], "%ogwart%", highlight: [title: p.title])
...> |> Repo.all()
...> |> Enum.map(& &1.title)
["H<b>ogwart</b>s Shocker"]Warning
The returned text is not HTML-escaped - the column content is returned
as-is with the matches wrapped in :start_sel/:stop_sel. If you render it as
raw HTML, either sanitize the result or use unique markers (e.g.
start_sel: "@@", stop_sel: "@@"), escape the result, and only then convert
the markers to tags.
Options
:type-:word(default) or:substring, see above.:start_sel,:stop_sel- strings the matches are wrapped in. Default to"<b>"and"</b>".
Options for type: :substring:
:case_sensitive- defaults tofalse(matchingilike/5); set totrueto only highlight exact-case occurrences (matchinglike/5).
Options for type: :word:
:language- language used for the search. Defaults to"english".:term_function- function used to convert the term tots_query. Same options as infull_text/5. Defaults to:websearch_to_tsquery.:prefix_search- whether to also highlight words the term matches as a prefix. Defaults totrue(same asfull_text/5).:highlight_all- whether to return the whole document.true(default) - returns the full text with all matches highlighted.false- returns a fragment (snippet) around the matches, controlled by the options below.
:max_words,:min_words- fragment size whenhighlight_all: false. Default to PostgreSQL's35and15.:short_word- words of this length or less are dropped at fragment start/end. Defaults to3.:max_fragments- maximum number of fragments to return. Defaults to0.:fragment_delimiter- string used to join fragments. Defaults to" ... ".
Examples
iex> insert_post!(title: "Hogwarts Shocker", body: "A spell disrupts the Quidditch Cup.")
...> Post
...> |> Torus.full_text([p], [p.title, p.body], "shocker")
...> |> select([p], Torus.highlight(p.title, "shocker"))
...> |> Repo.all()
["Hogwarts <b>Shocker</b>"]With a snippet instead of the full text:
iex> insert_post!(body: "A magic spell disrupts the Quidditch Cup final at Hogwarts.")
...> Post
...> |> Torus.full_text([p], [p.body], "quidditch")
...> |> select([p], Torus.highlight(p.body, "quidditch", highlight_all: false, max_words: 5, min_words: 2))
...> |> Repo.all()
["<b>Quidditch</b> Cup final"]Substring highlighting alongside ilike/5:
iex> insert_post!(title: "Hogwarts Shocker")
...> Post
...> |> Torus.ilike([p], [p.title], "%ogwart%")
...> |> select([p], Torus.highlight(p.title, "ogwart", type: :substring))
...> |> Repo.all()
["H<b>ogwart</b>s Shocker"]
Hybrid
Hybrid search: fuses several search strategies into a single ranked query using
Reciprocal Rank Fusion
(RRF). Each branch runs as an independent ranked subquery, keeps its limit best rows,
and the results are merged by summing weight * 1.0 / (k + rank) per row across branches.
Rows that rank high in several branches win; rows found by only one branch still
compete. This is the standard way to combine keyword (full_text/5, bm25/5) and
semantic (semantic/5) search, and generally outperforms each on its own.
Search branches
The third argument is a keyword list of search branches. Keys are search types -
:full_text, :similarity, :semantic, or :bm25 (pattern-match searches have no
ranking, so they can't participate). The same type can appear more than once. Values
are {qualifiers, term} or {qualifiers, term, opts} tuples mirroring the
corresponding search macro's arguments.
Branch opts accept the search type's own options (except :order, :score_key,
and :distance_key - branches are always ranked best-first, and the fused score is
exposed by hybrid/4 itself), plus:
:weight- multiplier for this branch's RRF score. Defaults to1.0.:limit- how many top rows this branch contributes. Defaults to20.:highlight- a keyword list of result keys to columns to highlight this branch's term matches in, e.g.highlight: [title: p.title]. Not supported in:semanticbranches. Seehighlight/3.
An empty search term contributes no rows to the fusion: full_text branches default empty_return to false, and
similarity/bm25 branches filter empty terms out.
Options
:k- RRF smoothing constant. Higher values flatten the difference between ranks. Defaults to60.:limit- final limit applied to the fused result.:score_key- atom key to select the fused score into the result map. The score is merged viaselect_merge/3, so the query needs to select a map before callinghybrid/4. The score is also available directly through the:torus_hybridnamed binding.:primary_key- column used to match rows across branches. Defaults to the schema's primary key.
Examples
iex> insert_post!(title: "Hogwarts Shocker", body: "A spell disrupts the Quidditch Cup.")
...> insert_post!(title: "Diagon Bombshell", body: "Secrets uncovered in the heart of Hogwarts.")
...> insert_post!(title: "Completely unrelated", body: "No magic here!")
...> Post
...> |> Torus.hybrid([p], [
...> full_text: {[p.title, p.body], "uncov hogwar"},
...> similarity: {[p.title], "hogwarts"}
...> ])
...> |> select([p], p.title)
...> |> Repo.all()
["Diagon Bombshell", "Hogwarts Shocker", "Completely unrelated"]With semantic search, weights, and the fused score selected:
search_vector = Torus.to_vector("A magic school in the UK")
Post
|> select([p], %{title: p.title})
|> Torus.hybrid([p], [
full_text: {[p.title, p.body], "magic school", weight: 1.0},
semantic: {p.embedding, search_vector, distance: :cosine_distance, weight: 2.0}
],
limit: 10,
score_key: :score
)
|> Repo.all()
# => [%{title: "...", score: 0.047}, ...]The fused query is a regular Ecto query - you can keep piping select, preload,
where, or pagination onto it. The base query's filters (everything piped in before
hybrid/4) apply to every branch. An order_by piped in before hybrid/4 is
discarded - the fused score defines the order - and a preload or offset piped in
before applies only to the fused result, not to the branches.
Optimizations
- Each branch is a separate subquery, so index each branch's search the same way
you would index the standalone search macro (GIN for full text and trigrams, HNSW /
IVFFlat for vectors, BM25 index for
bm25/5). - Branch
:limitcaps how many rows each branch ranks and contributes - keep it close to your final:limit(2x is a good default) so branches stay top-k friendly. - A branch without a filter (for example
similaritywithoutpre_filter: true, orfull_textwithfilter_type: :none) ranks every row the base query allows, which is a full scan without a matching index. Prefer filtered branches on large tables. :krarely needs tuning - 60 is the standard from the RRF paper and works well.
Pattern matching
Case-insensitive pattern matching search using
PostgreSQL ILIKE operator.
Warning
Doesn't clean the term, so it needs to be sanitized before being passed in. See
LIKE-injections.
You can use Torus.sanitize/1 to clean the term.
Options
:highlight- a keyword list of result keys to columns to highlight the term's matches in, e.g.highlight: [title: p.title]. Seehighlight/3.
Examples
iex> insert_posts!(titles: ["Wand", "Magic wand", "Owl"])
...> Post
...> |> Torus.ilike([p], [p.title], "wan%")
...> |> select([p], p.title)
...> |> Repo.all()
["Wand"]
iex> insert_posts!([%{title: "hogwarts", body: nil}, %{title: nil, body: "HOGWARTS"}])
...> Post
...> |> Torus.ilike([p], [p.title, p.body], "%OGWART%")
...> |> select([p], %{title: p.title, body: p.body})
...> |> order_by(:id)
...> |> Repo.all()
[%{title: "hogwarts", body: nil}, %{title: nil, body: "HOGWARTS"}]
iex> insert_post!(title: "MaGiC")
...> Post
...> |> Torus.ilike([p], p.title, "magi%")
...> |> select([p], p.title)
...> |> Repo.all()
["MaGiC"]Optimizations
See like/5 optimization section for more details.
Case-sensitive pattern matching search using PostgreSQL LIKE operator.
Warning
Doesn't clean the term, so it needs to be sanitized before being passed in. See
LIKE-injections.
You can use Torus.sanitize/1 to clean the term.
Options
:highlight- a keyword list of result keys to columns to highlight the term's matches in, e.g.highlight: [title: p.title]. Seehighlight/3.
Examples
iex> insert_posts!([%{title: "hogwarts", body: nil}, %{title: nil, body: "HOGWARTS"}])
...> Post
...> |> Torus.like([p], [p.title, p.body], "%OGWART%")
...> |> select([p], p.body)
...> |> Repo.all()
["HOGWARTS"]Optimizations
like/5is case-sensitive, so it can take advantage of B-tree indexes when there is no wildcard (%) at the beginning of the search term, prefer it overilike/5if possible.Adding a B-tree index:
CREATE INDEX index_posts_on_title ON posts (title);Use
GINorGiSTIndex withpg_trgmextension for LIKE and ILIKE.When searching for substrings (%word%), B-tree indexes won't help. Instead, use trigram indexing (
pg_trgmextension):CREATE EXTENSION IF NOT EXISTS pg_trgm; CREATE INDEX posts_title_trgm_idx ON posts USING GIN (title gin_trgm_ops);If using prefix search, convert data to lowercase and use B-tree index for case-insensitive search:
ALTER TABLE posts ADD COLUMN title_lower TEXT GENERATED ALWAYS AS (LOWER(title)) STORED; CREATE INDEX index_posts_on_title ON posts (title_lower);Torus.like([p], [p.title_lower], "hogwarts%")Use full-text search for large text fields, see
full_text/5for more details.
Removes all like/ilike special characters from the term, so it can be used in further pattern-match searches.
Examples
iex> Torus.sanitize(~S"%_\realterm%")
"realterm"
Similar to like/5, except that it interprets the pattern using the SQL standard's
definition of a regular expression. SQL regular expressions are a curious cross between
LIKE notation and common (POSIX) regular expression notation. See
PostgreSQL SIMILAR TO
Examples
iex> insert_post!(body: "abc")
...> Post
...> |> Torus.similar_to([p], [p.title, p.body], "%(b|d)%")
...> |> select([p], p.body)
...> |> Repo.all()
["abc"]Optimizations
- If regex is needed, use POSIX regex with
~or~*operators since they may leverage GIN or GiST indexes in some cases. These operators will be introduced later on. - Use
ilike/5orlike/5when possible,similar_to/5almost always does full table scans - Filter and limit the result set as much as possible before calling
similar_to/5
The substring function with three parameters provides extraction of a substring that matches an SQL regular expression pattern. The function can be written according to standard SQL syntax:
substring('foobar' similar '%#"o_b#"%' escape '#') oob
substring('foobar' similar '#"o_b#"%' escape '#') NULLExamples
insert_post!(title: "Hello123World")
Post |> select([p], substring(p.title, "[0-9]+", "#")) |> Repo.all()
["123"]
Semantic
Calls the specified embedding module's embedding_model/1 function to retrieve the model name.
See Semantic search guide for more info.
Semantic search using pgvector extension to compare vectors. See Semantic search guide for more info.
Options
:distance- a way to calculate the distance between the vectors. Can be one of::l2_distance(default) - L2 distance:max_inner_product- negative inner product:cosine_distance- cosine distance:l1_distance- L1 distance:hamming_distance- (binary vectors only) Hamming distance:jaccard_distance- (binary vectors only) Jaccard distance
:order- describes the ordering of the results. Possible values are:asc(default) - orders the results by distance in ascending order. 0 distance means that the vectors are the same, meaning the terms are equal. The closer the vectors - the more aligned are the terms.:desc- orders the results by distance in descending order.:none- doesn't apply ordering at all.
:pre_filter- a positive float that is passed directly to the query to pre-filter the results.:none(default) - no pre-filtering is done.float- pre-filters the results before applying the order. The results with vectors distance below the pre-filter value are returned.
:distance_key- pass an atom to put the selected distance under in the result map. The distance is merged viaselect_merge/3, so the query needs to select a map before callingsemantic/5(e.g.select([p], %{title: p.title})).:none(default) - the distance is not selected.atom- the map key the distance is put under.
Examples
def search(term) do
search_vector = Torus.to_vector(term)
Post
|> Torus.semantic([p], p.embedding, search_vector)
|> Repo.all()
endOptimizations
Use
pre_filterto pre-filter the results before applying the order. This would significantly reduce the number of rows to order.Index embeddings column:
HNSW (Hierarchical Navigable Small World) - High-accuracy Approximate Nearest Neighbor
CREATE INDEX ON embeddings USING hnsw (embedding vector_l2_ops) WITH (m = 16, ef_construction = 200);IVFFlat Index (Approximate Nearest Neighbor) with different similarity functions. Prior to index creation, it's recommended to have some real data in place, so the quality of clusters is better.
- Cosine Similarity
CREATE INDEX ON embeddings USING ivfflat (embedding vector_cosine_ops); - L2 Distance
CREATE INDEX ON embeddings USING ivfflat (embedding vector_l2_ops); - Inner Product
CREATE INDEX ON embeddings USING ivfflat (embedding vector_ip_ops);
- Cosine Similarity
Same as to_vectors/2, but returns the first vector from the list.
Takes a list of terms (binaries) and embedding module's specific options and passes them to embedding_module generate/2 function.
Configure embedding_module either in config.exs:
config :torus, :embedding_module, Torus.Embeddings.HuggingFaceor pass embedding_module as an option to to_vectors/2 function. Options always
have greater priority than the config.
See Semantic search guide for more info.
Similarity
Searches for records that closely match the input text using trigram distance. Ideal for fuzzy matching and catching typos in short text fields.
Implemented using case-insensitive similarity search using PostgreSQL similarity functions.
Warning
You need to have pg_trgm extension installed.
defmodule YourApp.Repo.Migrations.CreatePgTrgmExtension do
use Ecto.Migration
def change do
execute "CREATE EXTENSION IF NOT EXISTS pg_trgm", "DROP EXTENSION IF EXISTS pg_trgm"
end
endOptions
:type- similarity type. Possible options are::word_similarity(default) - usespg_trgmword_similarityfunction. Use it if you're dealing with sentences and you don't want the length of the strings to affect the search result.:strict_word_similarity- usesstrict_word_similarityfunction. Prioritizes full matches, forces extent boundaries to match word boundaries. Since we don't have cross-word trigrams, this function actually returns greatest similarity between first string and any continuous extent of words of the second string.:similarity- usessimilarityfunction. Compares the whole set of trigrams instead of an ordered subset. Use this when you search for the exact phrase or a string, not a word/phrase in a sentence/longer text.
:order- describes the ordering of the results. Possible values are:desc(default) - orders the results by similarity rank in descending order.:asc- orders the results by similarity rank in ascending order.:none- doesn't apply ordering at all.
:pre_filter- whether or not to pre-filter the results:false(default) - omits pre-filtering and returns all results.true- before applying the order, pre filters (using boolean operators which potentially use GIN indexes) the result set. The results abovepg_trgm.{type}_thresholdare returned. It is advised to set the corresponding value to0.3so that more relevant results are returned. For example, forword_similarity, we'd runSET pg_trgm.word_similarity_threshold = 0.3;.
:highlight- a keyword list of result keys to columns to highlight the term's exact-word matches in, e.g.highlight: [title: p.title]. Seehighlight/3.
Examples
iex> insert_post!(title: "Hogwarts Shocker", body: "A spell disrupts the Quidditch Cup.")
...> insert_post!(title: "Diagon Bombshell", body: "Secrets uncovered in the heart of Hogwarts.")
...> insert_post!(title: "Completely unrelated", body: "No magic here!")
...> Post
...> |> Torus.similarity([p], [p.title, p.body], "Diagon Bombshell")
...> |> limit(1)
...> |> select([p], p.title)
...> |> Repo.all()
["Diagon Bombshell"]
iex> insert_posts!(["Wand", "Owl", "What an amazing cloak"])
...> Post
...> |> Torus.similarity([p], [p.title], "owls", pre_filter: true)
...> |> select([p], p.title)
...> |> Repo.all()
["Owl"]Optimizations
- Use
pre_filter: trueto pre-filter the results before applying the order. This would significantly reduce the number of rows to order. The pre-filtering phase uses different (boolean) similarity operators which more actively leverage GIN indexes. - Use
order: :noneargument if you don't care about the order of the results. The query will return all results that are above the similarity threshold, which you can set globally viaSET pg_trgm.{type}_threshold = 0.3;, replacing type with yourtypeoption (e.g.word_similarity_threshold). - When
order: :desc(default) and the limit is not set, the query will do a full table scan, so it's recommended to manually limit the results (by applyingwhereorlimitclauses to filter the rows as much as possible).
Adding an index
-- If you haven't created it yet
CREATE EXTENSION IF NOT EXISTS pg_trgm;
CREATE INDEX index_posts_on_title ON posts USING GIN (title gin_trgm_ops);