Repository navigation
Update Elastic Search, QDrant, and PGVector to retrieve_online_documents_v2 method #5115
Description
Activity
- added a parent issue
on Mar 4, 2025 @franciscojavierarceo can I work on this issue ?
Reacted by Francisco Javier Arceofranciscojavierarceo commented
on Mar 31, 2025 MemberAuthorMore actionsOf course!
Reacted by Yassin Nouh@franciscojavierarceo can I create a pr for each one starting with Qdrant ?
franciscojavierarceo commented
on Apr 1, 2025 MemberAuthorMore actionsYeah of course!
For the implementation of
retrieve_online_documents_v2with PGVector, I compared the available options for text search.There are two main approaches:
-
Using
LIKEorILIKE- Simple pattern matching for substrings.
- Easy to implement, but does not support ranking, stemming, or advanced text analysis.
- Performance is not optimal for large datasets, especially without proper indexing.
- Does not integrate well if we plan to combine text and vector scores.
-
Using PostgreSQL Full-Text Search (
to_tsvectorandto_tsquery)- Tokenizes text, removes stop words, and supports stemming.
- Supports relevance ranking.
- Performs better on larger datasets when combined with GIN indexing.
- Can work in combination with vector search for hybrid retrieval.
- Example:
WHERE to_tsvector('english', content) @@ to_tsquery('english', 'keyword')
Recommendation:
to_tsvectorandto_tsqueryfull-text search is more effective based on my view. but we need to change PGvector config as it will require languageReacted by Francisco Javier Arceo-
franciscojavierarceo commented
on Apr 7, 2025 MemberAuthorMore actionsNice, I think (2) sounds good. I'd probably mention (1) in another ticket in case someone wants to pick it up in the future or compare it.
Reacted by Yassin NouhI noticed that in the v2 method there isn’t any configuration option for specifying the language for Lemxme. Without this setting, we might run into limitations when handling non-English content, especially since full-text search features (like stemming and tokenization in PostgreSQL's to_tsvector/to_tsquery) can benefit from language-specific configuration. It might be worth exploring the addition of a language parameter to make the method more flexible and accurate for diverse datasets. What do you think?
franciscojavierarceo commented
on Apr 7, 2025 MemberAuthorMore actionsYeah that makes a lot of sense! Do you mind defaulting it to
en-us?Reacted by Yassin Nouhit will look like this
query = sql.SQL( """ SELECT entity_key, feature_name, value, vector_value, NULL as distance, ts_rank(to_tsvector('english', value_text), to_tsquery('english', %s)) as text_rank, event_ts, created_ts FROM {table_name} WHERE feature_name = ANY(%s) AND to_tsvector('english', value_text) @@ to_tsquery('english', %s) ORDER BY text_rank DESC LIMIT {top_k} """ ).format( table_name=sql.Identifier(table_name), top_k=sql.Literal(top_k), )
franciscojavierarceo commented
on Apr 7, 2025 MemberAuthorMore actionshmm, we could probably have a dict for that lookup.
iso_lan_dict = { "en": "english", "es": "spanish", } Or something around there. The point being to map to [ISO language codes](https://en.wikipedia.org/wiki/List_of_ISO_639_language_codes)
Reacted by Yassin Nouh
Is your feature request related to a problem? Please describe.
Update Elastic Search, QDrant, and PGVector to
retrieve_online_documents_v2method, as it returns broader features.Describe the solution you'd like
V2 is the better method.
Describe alternatives you've considered
N/A
Additional context
N/A