Visitar URL original
Historical retrieval without an entity dataframe · Issue #1611 · feast-dev/feast · GitHub
Skip to content

Historical retrieval without an entity dataframe #1611

Description

@woop

Is your feature request related to a problem? Please describe.
The current Feast get_historical_features() method requires that users provide an entity dataframe as follows

training_df = store.get_historical_features(
    entity_df=entity_df, 
    feature_refs = [
        'drivers_activity:trips_today'
        'drivers_profile:rating'
    ],
)

However, many users would like the feature store to provide entities to them for training, instead of having to query or provide entities as part of the entity dataframe.

Describe the solution you'd like
Allow users to specify an existing feature view from which an entity dataframe will be queried.

training_df = store.get_historical_features(
    entity_df="drivers_activity", 
    feature_refs = [
        'drivers_activity:trips_today'
        'drivers_profile:rating'
    ],
)

With the addition of time range filtering.

training_df = store.get_historical_features(
    entity_df="drivers_activity", 
    feature_refs = [
        'drivers_activity:trips_today'
        'drivers_profile:rating'
    ],
    from_date=(today - timedelta(days = 7)),
    to_date=datetime.now(),
)

Activity

  1. tedhtchang commented on Jun 7, 2021

    @tedhtchang
    Contributor

    training_df = store.get_historical_features(
    left_table="drivers_activity",
    feature_refs = [
    'drivers_activity:trips_today'
    'drivers_activity:rating'
    ],
    )

    Does this mean the resulting training_df contain every row (but only selectdriver_id, event_timestamp, trips_today, and rating columns), from the drivers_activity view ?

  2. woop commented on Jun 7, 2021

    @woop
    MemberAuthor

    training_df = store.get_historical_features(
    left_table="drivers_activity",
    feature_refs = [
    'drivers_activity:trips_today'
    'drivers_activity:rating'
    ],
    )

    Does this mean the resulting training_df contain every row (but only selectdriver_id, event_timestamp, trips_today, and rating columns), from the drivers_activity view ?

    Actually my example was poor. I've modified it to show that we can query multiple feature views. Essentially how it works is that we will query the entity_df for all entities, but it can now be an existing feature view. We would only query it for timestamps and entity columns. Features then get joined onto those rows as usual.

  3. NoRaincheck commented on Jul 2, 2021

    @NoRaincheck
    Contributor

    Should there also be an option to "keep latest" only, when used in conjunction with the time range filtering
    Otherwise its more than possible that the underlying entity dataframe could have duplicated entity keys.

    The usecase for this in my mind is for backtesting purposes.

  4. woop commented on Jul 2, 2021

    @woop
    MemberAuthor

    Should there also be an option to "keep latest" only, when used in conjunction with the time range filtering
    Otherwise its more than possible that the underlying entity dataframe could have duplicated entity keys.

    The usecase for this in my mind is for backtesting purposes.

    Do you mean entity row or entity key? https://docs.feast.dev/concepts/data-model-and-concepts#entity-row

    So you would not want to return features with the same entity key over different dates?

  5. NoRaincheck commented on Jul 2, 2021

    @NoRaincheck
    Contributor

    I was thinking entity key. Only as an option - there are use cases for enabling both of them.

    For example, if our machine learning deployment is a daily batch job, perhaps for back-testing we would have the get_historical_features(from_date=my_date-timedelta(days=1), to_date=my_date), where my_date is the timestamp to simulate when our machine learning job "would run" on a daily basis

    Though this then raises a good question on how this kind of workflow should be productionised? E.g. if I have an hourly/daily batch which goes through our whole customer base to find fraudulent customers, how should this work in feast? We wouldn't really use the online store for this, and this API could look something like:

    my_daily_batch_scoring_df = store.get_historical_features(
        entity_df = "my_df", 
        feature_refs = [...],
        latest=True,
        from_date=(today - timedelta(days = 1)),
        to_date=datetime.now(),
    )
    

    Probably a discussion for another thread...

  6. NoRaincheck commented on Jul 2, 2021

    @NoRaincheck
    Contributor

    Can I give this a go and raise a PR for File based offlinestore only?

    I'll stick the spec written(?), though I noticed elsewhere in the repo the nomenclature used was start_date and end_date - should we align to that rather than from_date and to_date https://github.com/feast-dev/feast/blob/master/sdk/python/feast/infra/offline_stores/file.py#L219-L220 ?

  7. woop commented on Jul 5, 2021

    @woop
    MemberAuthor

    I was thinking entity key. Only as an option - there are use cases for enabling both of them.

    For example, if our machine learning deployment is a daily batch job, perhaps for back-testing we would have the get_historical_features(from_date=my_date-timedelta(days=1), to_date=my_date), where my_date is the timestamp to simulate when our machine learning job "would run" on a daily basis

    Though this then raises a good question on how this kind of workflow should be productionised? E.g. if I have an hourly/daily batch which goes through our whole customer base to find fraudulent customers, how should this work in feast? We wouldn't really use the online store for this, and this API could look something like:

    my_daily_batch_scoring_df = store.get_historical_features(
        entity_df = "my_df", 
        feature_refs = [...],
        latest=True,
        from_date=(today - timedelta(days = 1)),
        to_date=datetime.now(),
    )
    

    Probably a discussion for another thread...

    I can see the value in this. In fact, some other folks have also asked for it. Would you mind creating a new issue and linking back to this issue for us? I think it's worth a separate discussion. Specifically, the need for a latest only argument in get_historical_features().

  8. woop commented on Jul 5, 2021

    @woop
    MemberAuthor

    Can I give this a go and raise a PR for File based offlinestore only?

    I'll stick the spec written(?), though I noticed elsewhere in the repo the nomenclature used was start_date and end_date - should we align to that rather than from_date and to_date https://github.com/feast-dev/feast/blob/master/sdk/python/feast/infra/offline_stores/file.py#L219-L220 ?

    You can give it a go, but we probably won't release it until we have support for all our main stores. Perhaps a better middle ground is to add a new method to the FeatureStore class and have it throw a NotImplemented exception for the other stores, and specifically print warnings that this functionality is experimental and will change.

  9. NoRaincheck commented on Jul 5, 2021

    @NoRaincheck
    Contributor

    Sounds good, hopefully I'll pull something together "soon". I'll name the method something sensible as well.

  10. MattDelac commented on Jul 15, 2021

    @MattDelac
    Collaborator

    Hi there 👋 ,

    As I already explained to Willem, we built an higher level API on our side to make the life of our users easier

    It basically does the following

    def get_historical_features(
        feature_refs: List[str],
        threshold: Union[datetime, date] = None,
        sample_size: int = 1000,
        left_feature_view: Union[pd.DataFrame, str] = None,
        full_feature_names: bool = False,
    ) -> BigQueryRetrievalJob:
        # If all the features come from the same FeatureView then we infer the `left_feature_view` parameter
    
        # We get the unique_join_keys in order to remove some duplicate data if it exists
        # It's more or less the following
        query = f"""
            SELECT
                {', '.join(unique_join_keys)},
                TIMESTAMP '{str_timestamp}' AS {timestamp_column}
            FROM {source_table}
            {where_clause}
            GROUP BY {', '.join(unique_join_keys)}
            {limit_clause}
        """
        # The limit_clause only exist if we want a sample of the left FeatureView
    
        store = FeatureStore()
        # We build the query for our users and pass it to Feast
        return store.get_historical_features(
            entity_df=sql_query,
            feature_refs=features,
            full_feature_names=full_feature_names,
        )

    Happy to have a chat about a similar API implemented in Feast

  11. stale commented on Nov 14, 2021

    @stale

    This issue has been automatically marked as stale because it has not had recent activity. It will be closed if no further activity occurs. Thank you for your contributions.

  12. added and removed
    wontfixThis will not be worked on
    on Nov 21, 2021
  13. 18 remaining items

  14. franciscojavierarceo commented on Jan 9, 2026

    @franciscojavierarceo
    Member

    @jyejare this is great, can you make child issues for all of the other offline stores?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions