Visitar URL original
Update Elastic Search, QDrant, and PGVector to `retrieve_online_documents_v2` method · Issue #5115 · feast-dev/feast · GitHub
Skip to content

Update Elastic Search, QDrant, and PGVector to retrieve_online_documents_v2 method #5115

Description

@franciscojavierarceo

Is your feature request related to a problem? Please describe.
Update Elastic Search, QDrant, and PGVector to retrieve_online_documents_v2 method, as it returns broader features.

Describe the solution you'd like
V2 is the better method.

Describe alternatives you've considered
N/A

Additional context
N/A

Activity

  1. YassinNouh21 commented on Mar 31, 2025

    @YassinNouh21
    Collaborator

    @franciscojavierarceo can I work on this issue ?

  2. franciscojavierarceo commented on Mar 31, 2025

    @franciscojavierarceo
    MemberAuthor

    Of course!

  3. YassinNouh21 commented on Apr 1, 2025

    @YassinNouh21
    Collaborator

    @franciscojavierarceo can I create a pr for each one starting with Qdrant ?

  4. franciscojavierarceo commented on Apr 1, 2025

    @franciscojavierarceo
    MemberAuthor

    Yeah of course!

  5. YassinNouh21 commented on Apr 6, 2025

    @YassinNouh21
    Collaborator

    @franciscojavierarceo ,

    For the implementation of retrieve_online_documents_v2 with PGVector, I compared the available options for text search.

    There are two main approaches:

    1. Using LIKE or ILIKE

      • Simple pattern matching for substrings.
      • Easy to implement, but does not support ranking, stemming, or advanced text analysis.
      • Performance is not optimal for large datasets, especially without proper indexing.
      • Does not integrate well if we plan to combine text and vector scores.
    2. Using PostgreSQL Full-Text Search (to_tsvector and to_tsquery)

      • Tokenizes text, removes stop words, and supports stemming.
      • Supports relevance ranking.
      • Performs better on larger datasets when combined with GIN indexing.
      • Can work in combination with vector search for hybrid retrieval.
      • Example:
        WHERE to_tsvector('english', content) @@ to_tsquery('english', 'keyword')

    Recommendation: to_tsvector and to_tsquery full-text search is more effective based on my view. but we need to change PGvector config as it will require language

  6. franciscojavierarceo commented on Apr 7, 2025

    @franciscojavierarceo
    MemberAuthor

    Nice, I think (2) sounds good. I'd probably mention (1) in another ticket in case someone wants to pick it up in the future or compare it.

  7. YassinNouh21 commented on Apr 7, 2025

    @YassinNouh21
    Collaborator

    @franciscojavierarceo

    I noticed that in the v2 method there isn’t any configuration option for specifying the language for Lemxme. Without this setting, we might run into limitations when handling non-English content, especially since full-text search features (like stemming and tokenization in PostgreSQL's to_tsvector/to_tsquery) can benefit from language-specific configuration. It might be worth exploring the addition of a language parameter to make the method more flexible and accurate for diverse datasets. What do you think?

  8. franciscojavierarceo commented on Apr 7, 2025

    @franciscojavierarceo
    MemberAuthor

    Yeah that makes a lot of sense! Do you mind defaulting it to en-us?

  9. YassinNouh21 commented on Apr 7, 2025

    @YassinNouh21
    Collaborator

    it will look like this

        query = sql.SQL(
                        """
                        SELECT 
                            entity_key,
                            feature_name,
                            value,
                            vector_value,
                            NULL as distance,
                            ts_rank(to_tsvector('english', value_text), to_tsquery('english', %s)) as text_rank,
                            event_ts,
                            created_ts
                        FROM {table_name}
                        WHERE feature_name = ANY(%s) AND to_tsvector('english', value_text) @@ to_tsquery('english', %s)
                        ORDER BY text_rank DESC
                        LIMIT {top_k}
                        """
                    ).format(
                        table_name=sql.Identifier(table_name),
                        top_k=sql.Literal(top_k),
                    )
  10. franciscojavierarceo commented on Apr 7, 2025

    @franciscojavierarceo
    MemberAuthor

    hmm, we could probably have a dict for that lookup.

    iso_lan_dict = {
        "en": "english",
        "es": "spanish",
    }
    
    Or something around there. The point being to map to [ISO language codes](https://en.wikipedia.org/wiki/List_of_ISO_639_language_codes)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions