Semantic Search

Semantic Search in Siren Investigate allows you to search your data based on meaning rather than exact keyword matches. It supports two field types: semantic_text and dense_vector. Both types integrate with the:

  • dashboards where users can create a semantic filter using the filter editor

  • global search via dynamic filters

Semantic Text

The semantic_text field type uses an inference endpoint to automatically generate text embeddings. When you search on a semantic_text field, Elasticsearch runs the query through the configured inference model and returns results that are semantically similar.

semantic_text support in Investigate requires Elasticsearch 8.19.14 or later, an Elasticsearch enterprise license, and Siren Federate 8.19.14-39.2 or later. If the semantic_text type is unavailable in your Elasticsearch cluster, any dynamic filter mapped to a semantic_text field will be disabled with an explanatory warning.

In order to index a document containing a semantic_text field, the user’s role must have the indices:admin/mapping/auto_put permission.

semantic_text search relies on Elasticsearch’s machine learning features (xpack.ml), which are not available on Intel-based macOS machines. On Intel Macs, any dynamic filter mapped to a semantic_text field will be disabled with an explanatory warning.

Due to an Elasticsearch design decision, the _field_caps API reports semantic_text fields as text. Siren Investigate automatically cross-references the _mapping API to detect the real field type and handles semantic_text fields correctly throughout the application.

Adding a semantic_text field in the data model

  1. In the Data model app, select the entity table you want to add the field to and navigate to the fields tab.

  2. Click Edit Mode then Add Fields and set the field type to Semantic Text.

  3. Optionally, in the Advanced mapping section, specify the inference endpoint to use for generating embeddings. If omitted, the cluster default ELSER endpoint is used automatically.

  4. Click Save.

Using semantic_text in dashboard filters

semantic_text fields support the following filter operators in the filter editor:

  • is / is not: Uses a match query against the field.

  • is one of / is not one of: Uses a matches query against the field.

Standard phrase and term filter operators are not supported for semantic_text fields.

Using semantic_text dynamic filters

  1. In the Data model app, click Dynamic filters and click Create dynamic filter.

  2. Set the filter type to semantic_text.

  3. Click Save.

  4. Now, select the entity table and go to the Dynamic filters tab.

  5. Click Add dynamic filters, select the filter you created, and map it to the field of type semantic_text that you want to search on.

  6. Click Save.

Once the filter is mapped, it appears in the global search when searching that entity table, and can also be used as a dashboard filter. Enter any natural-language text in the filter input field and Siren Investigate will pass the query to the inference endpoint and return semantically matching records.

Known limitations

  • When using a third-party inference model (for example, OpenAI) with a match query with size = 0, the hits.total.value in the Elasticsearch response is inaccurate and always returns k (nearest neighbors) * the number of should clauses, even when track_total_hits: true is set.

Dense Vector

The dense_vector field type stores fixed-length arrays of floating-point numbers and enables vector similarity search using pre-computed embeddings. When you search a dense_vector field, Siren Investigate sends your input text to the configured embedding model, converts it to a vector, and performs a k-nearest neighbour (kNN) search against the stored vectors to return the most similar records.

dense_vector support in Investigate requires Elasticsearch 8.15.5 or later and Siren Federate 8.15.5-37.2 or later. Additionally, the Siren AI plugin must be installed and configured with at least one text embedding model. If the embeddings API is not configured, any dynamic filter mapped to a dense_vector field will be disabled with an explanatory warning.

Adding a dense_vector field in the data model

  1. In the Data model app, select the entity table you want to add the field to and navigate to the fields tab.

  2. Click Edit Mode then Add Fields and set the field type to Dense Vector.

  3. Optionally, in the Advanced mapping section, specify the vector properties in JSON syntax:

    {
      "meta": { "model": "text-embedding-3-small" },
      "type": "dense_vector",
      "dims": 1536
    }
    Property Description

    meta.model

    The name of the embedding model to use for this field. Must correspond to a model configured in the Siren AI plugin.

    type

    Must be dense_vector.

    dims

    The number of dimensions in each embedding vector. Must match the output dimensionality of your embedding model.

  4. Click Save.

The specific embedding model used at index time must be defined in the meta.model field of the mapping. This ensures that the same model is used consistently for both ingestion and querying.

If a reindex operation is performed with a different model, or with a change to the dims parameter, the model configuration on the entity table must also be updated. This can be done via a fields refresh on the entity table.

Configuring the embeddings API

The Siren AI plugin must be enabled and have at least one text embedding model configured (embed:text capability) for dense vector search to work.

Instructions for installing the Siren AI plugin can be found here.

Configuring the mapping between text and vector

When computing embeddings for larger amounts of text, the text must be split into smaller chunks due to embedding model limitations. To tell the system which dense_vector field should store automatically computed embeddings, go to Data model → Semantic search.

Currently, the following scenarios are supported:

No chunking: Use this when the source text contains small content (for example, messages or summaries) that does not need to be split and always produces a single vector.

For this scenario, select:

  • source text field

  • corresponding dense_vector field

Chunk to nested fields: Use this when the source text contains larger content (for example, descriptions or full documents) that must be split into smaller chunks and produces multiple vectors.

For this scenario, select:

  • source text field

  • corresponding nested dense_vector field

  • corresponding nested text chunk field (optional), where the chunk text fragment is stored

  • corresponding nested integer chunk number field (optional), where the chunk number is stored

  • corresponding nested integer chunk start field (optional), where the chunk start offset in characters is stored

  • corresponding nested integer chunk end field (optional), where the chunk end offset in characters is stored

This configuration allows the system to compute the corresponding dense_vector field and optional chunk metadata when:

  • creating or editing documents in Document Viewer

  • creating or editing an "is semantically similar" filter in the filter editor

  • applying a dynamic filter in Siren Search or Global Search

If the mapping between source text and dense_vector fields is not configured, some semantic search features may not work correctly.

Using dense_vector in dashboard filters

dense_vector fields support the following filter operators in the filter editor:

  • is / is not: Uses a vector similarity query against the field.

  • exists / does not exist: Checks for the presence of a vector value in the field.

Using dense_vector dynamic filters

  1. In the Data model app, click Dynamic filters and click Create dynamic filter.

  2. Set the filter type to Dense_vector(Text).

  3. Click Save.

  4. Now, select the entity table and go to the Dynamic filters tab.

  5. Click Add dynamic filters, select the filter you created, and map it to the field of type dense_vector that you want to search on.

  6. Click Save.

Once the filter is mapped, it appears in the global search when searching that entity table, and can also be used as a dashboard filter. Enter any natural-language text in the filter input field and Siren Investigate will pass the query to the inference endpoint and return semantically matching records.

Known limitations

  • When querying a dense_vector field with size: 0, the returned hit count is approximate and typically equals k × number_of_shards rather than the true document count.

  • Elasticsearch does not support highlighting on dense_vector fields. Highlight results will not be returned for these fields.