Semantic Search
Semantic Search in Siren Investigate allows you to search your data based on meaning rather than exact keyword matches.
It supports two field types: semantic_text and dense_vector. Both types integrate with the:
-
dashboards where users can create a semantic filter using the filter editor
-
global search via dynamic filters
Semantic Text
The semantic_text field type uses an inference endpoint to automatically generate text embeddings.
When you search on a semantic_text field, Elasticsearch runs the query through the configured inference model
and returns results that are semantically similar.
|
|
|
In order to index a document containing a |
|
|
|
Due to an Elasticsearch design decision,
the |
Adding a semantic_text field in the data model
-
In the Data model app, select the entity table you want to add the field to and navigate to the fields tab.
-
Click Edit Mode then Add Fields and set the field type to Semantic Text.
-
Optionally, in the Advanced mapping section, specify the inference endpoint to use for generating embeddings. If omitted, the cluster default ELSER endpoint is used automatically.
-
Click Save.
Using semantic_text in dashboard filters
semantic_text fields support the following filter operators in the filter editor:
-
is / is not: Uses a
matchquery against the field. -
is one of / is not one of: Uses a
matchesquery against the field.
|
Standard phrase and term filter operators are not supported for |
Using semantic_text dynamic filters
-
In the Data model app, click Dynamic filters and click Create dynamic filter.
-
Set the filter type to semantic_text.
-
Click Save.
-
Now, select the entity table and go to the Dynamic filters tab.
-
Click Add dynamic filters, select the filter you created, and map it to the field of type
semantic_textthat you want to search on. -
Click Save.
Once the filter is mapped, it appears in the global search when searching that entity table, and can also be used as a dashboard filter. Enter any natural-language text in the filter input field and Siren Investigate will pass the query to the inference endpoint and return semantically matching records.
Dense Vector
The dense_vector field type stores fixed-length arrays of floating-point numbers and enables
vector similarity search using pre-computed embeddings.
When you search a dense_vector field, Siren Investigate sends your input text to the configured
embedding model, converts it to a vector, and performs a k-nearest neighbour (kNN) search
against the stored vectors to return the most similar records.
|
|
Adding a dense_vector field in the data model
-
In the Data model app, select the entity table you want to add the field to and navigate to the fields tab.
-
Click Edit Mode then Add Fields and set the field type to Dense Vector.
-
Optionally, in the Advanced mapping section, specify the vector properties in JSON syntax:
{ "meta": { "model": "text-embedding-3-small" }, "type": "dense_vector", "dims": 1536 }Property Description meta.modelThe name of the embedding model to use for this field. Must correspond to a model configured in the Siren AI plugin.
typeMust be
dense_vector.dimsThe number of dimensions in each embedding vector. Must match the output dimensionality of your embedding model.
-
Click Save.
|
The specific embedding model used at index time must be defined in the |
|
If a reindex operation is performed with a different model, or with a change to the |
Configuring the embeddings API
The Siren AI plugin must be enabled and have at least one text embedding model configured
(embed:text capability) for dense vector search to work.
Instructions for installing the Siren AI plugin can be found here.
Configuring the mapping between text and vector
When computing embeddings for larger amounts of text, the text must be split into smaller chunks due to embedding model limitations. To tell the system which dense_vector field should store automatically computed embeddings, go to Data model → Semantic search.
Currently, the following scenarios are supported:
No chunking: Use this when the source text contains small content (for example, messages or summaries) that does not need to be split and always produces a single vector.
For this scenario, select:
-
source text field
-
corresponding dense_vector field
Chunk to nested fields: Use this when the source text contains larger content (for example, descriptions or full documents) that must be split into smaller chunks and produces multiple vectors.
For this scenario, select:
-
source text field
-
corresponding nested dense_vector field
-
corresponding nested text chunk field (optional), where the chunk text fragment is stored
-
corresponding nested integer chunk number field (optional), where the chunk number is stored
-
corresponding nested integer chunk start field (optional), where the chunk start offset in characters is stored
-
corresponding nested integer chunk end field (optional), where the chunk end offset in characters is stored
This configuration allows the system to compute the corresponding dense_vector field and optional chunk metadata when:
-
creating or editing documents in Document Viewer
-
creating or editing an "is semantically similar" filter in the filter editor
-
applying a dynamic filter in Siren Search or Global Search
If the mapping between source text and dense_vector fields is not configured, some semantic search features may not work correctly.
Using dense_vector in dashboard filters
dense_vector fields support the following filter operators in the filter editor:
-
is / is not: Uses a vector similarity query against the field.
-
exists / does not exist: Checks for the presence of a vector value in the field.
Using dense_vector dynamic filters
-
In the Data model app, click Dynamic filters and click Create dynamic filter.
-
Set the filter type to Dense_vector(Text).
-
Click Save.
-
Now, select the entity table and go to the Dynamic filters tab.
-
Click Add dynamic filters, select the filter you created, and map it to the field of type
dense_vectorthat you want to search on. -
Click Save.
Once the filter is mapped, it appears in the global search when searching that entity table, and can also be used as a dashboard filter. Enter any natural-language text in the filter input field and Siren Investigate will pass the query to the inference endpoint and return semantically matching records.
Known limitations
-
When querying a
dense_vectorfield withsize: 0, the returned hit count is approximate and typically equalsk × number_of_shardsrather than the true document count. -
Elasticsearch does not support highlighting on
dense_vectorfields. Highlight results will not be returned for these fields.