Visitar URL original
Improve Feature Server Observability · Issue #5920 · feast-dev/feast · GitHub
Skip to content

Improve Feature Server Observability #5920

Description

@franciscojavierarceo

Is your feature request related to a problem? Please describe.

The Feature Server currently has limited built-in observability. While there is initial OpenTelemetry (OTEL) support, it does not expose standard service-level RED metrics (Rate, Errors, Duration) needed to operate the server in production.

Users need visibility into throughput, error rates, and latency for the core APIs:

  • /get-online-features
  • /retrieve-online-documents
  • /push
  • /write-to-online-store
  • /materialize
  • /materialize-incremental

Today this often requires custom middleware or external tooling.

Describe the solution you'd like

Extend Feature Server’s OTEL integration to emit standard RED metrics out of the box, with useful Feast-specific breakdowns.

Metrics per endpoint should include:

  • Request rate/count
  • Error count/rate (by HTTP status class, optionally basic error categories)
  • Latency histograms (supporting p50/p95/p99)

Additional breakdowns:

  • Break out metrics by Feature Service name (and Feature View name where applicable/available), with safeguards to limit label cardinality (e.g., allowlist and/or config flags; default off if needed).
  • Segment latency by number of features requested using configurable bins (e.g., 1–10, 11–50, 51–200, 201+).

Implementation note (based on current feature_server.py structure):

  • Add a FastAPI middleware to record RED metrics for every request (keyed by endpoint + status_class).
  • Populate Feast-specific labels (feature_service / feature_view / feature_count_bin, etc.) inside the request handlers (since they come from the parsed request body) and attach them via request.state for the middleware to include.
  • For long-running endpoints (/materialize*), ensure duration histograms support multi-second/minute latencies.

Optionally expose basic internal timings via tracing spans and/or additional histograms:

  • online store read/write duration
  • on-demand transformation execution duration
  • materialization step duration

Metrics should use native OpenTelemetry APIs and work with common OTEL collectors/exporters (e.g., Prometheus, OpenTelemetry Collector, Grafana).

Describe alternatives you've considered

  • Custom HTTP middleware
  • Service mesh / proxy metrics (miss Feast-specific context like FS/FV/feature_count)
  • Manual instrumentation

Additional context

Standard RED metrics with Feast-aware labels would significantly improve Feature Server operability and align Feast with common production observability practices, while keeping label cardinality under control.

Activity

  1. anshishrivastava commented on Feb 5, 2026

    @anshishrivastava
    Contributor

    Hi @franciscojavierarceo — I'd like to own this issue.

    Before implementation, I'll prepare a lightweight design doc covering:

    Metrics Design

    • RED metrics per endpoint: request rate, error rate (by status code class), duration (p50/p95/p99)
    • Labeling strategy: feature_service, feature_view, entity_key_count with cardinality controls to prevent label explosion in high-throughput production environments
    • Histogram buckets: logarithmic bucketing for latency (e.g., 1ms-10s range) tailored to typical online serving SLOs
    • Error taxonomy: distinguish between client errors (invalid requests, missing features), server errors (registry/provider failures), and upstream errors (offline store timeouts)

    Implementation Considerations

    • OTel integration: native OpenTelemetry instrumentation with semantic conventions for observability, ensuring compatibility with Prometheus (via OTLP bridge), Datadog, and cloud-native collectors
    • Performance impact: async metric recording to avoid adding latency to the critical serving path; batch aggregation where applicable
    • Registry-aware labeling: dynamic label extraction from the feature store registry to maintain consistency with Feast's metadata model

    Proposed Implementation Plan

    1. Middleware instrumentation + handler-level metric emission
    2. Unit tests (metric correctness) + integration tests (end-to-end validation)
    3. Example Grafana dashboard + basic alerting rules + quickstart docs

    Happy to iterate on the design before implementation.
    Please assign if this approach aligns with the roadmap. Thanks!

  2. ntkathole commented on Feb 5, 2026

    @ntkathole
    Member

    Thank you @anshishrivastava for showing interest in it. Overall I see it aligns.
    Feel free to tag @jyejare if you have some questions or need any clarification.

  3. jyejare commented on Feb 6, 2026

    @jyejare
    Collaborator

    Hello @anshishrivastava , Thanks for showing your interest in the issue. Currently, we are in the process of developing RED metrics in the repo for OpenDataHub purpose; the PR will be up soon. But we are not covering all the metrics detailed in your design.

    Sorry to say this but, I would want you to hold until implementation is done from our side, and then build on top it later.

    Thanks for your understanding and patience. I am happy to answer any queries you have.

  4. anshishrivastava commented on Feb 25, 2026

    @anshishrivastava
    Contributor

    Hi @jyejare, checking in on the status of the OpenDataHub RED metrics PR.

    Has that work been merged yet? I'm still very interested in building out the remaining metrics and Feast-specific labels we discussed (like feature_service and feature_view segmentation) to complete the observability story.

    If the base PR is up, I can start reviewing it or prepare a follow-up PR on top of it. Let me know how you'd like me to proceed!

  5. jyejare commented on Feb 25, 2026

    @jyejare
    Collaborator

    @anshishrivastava Hey, thanks for the follow up. The work is postponed for a few days for other priorities, I shall keep you posted or at least I ll provide the top up work to you.

  6. ntkathole commented on Jul 21, 2026

    @ntkathole
    Member

    @jyejare can you please confirm if this is available now? Please also add If something can be still worked on

  7. jyejare commented on Jul 21, 2026

    @jyejare
    Collaborator

    @ntkathole @anshishrivastava — Here's an update on the current state.

    What's available now

    The core Feature Server observability is in place. The following have been merged:

    Capability PR Status
    Online store RED metrics (request rate, error rate, latency histograms per endpoint) Pre-existing in metrics.py ✅ Available
    Offline store RED metrics (feast_offline_store_request_total, _latency_seconds, _row_count) #6340 ✅ Available
    SOX audit logging (online + offline, structured JSON via feast.audit logger) #6340 ✅ Available
    Online store read/write duration histogram Pre-existing ✅ Available
    On-demand transformation duration histogram Pre-existing ✅ Available
    Materialization duration + result metrics Pre-existing ✅ Available
    Feature quality monitoring (DQM) — full backend, 8 offline stores, REST API, CLI, auto-baseline #6202 ✅ Available
    DQM UI (monitoring dashboard with histograms, time-series, filters) #6422 ✅ Available
    Batch + Log data source support for DQM #6202 ✅ Available

    What can still be worked on

    The following items from this issue are open for contribution:

    1. Native OpenTelemetry SDK migration — Current metrics use the Prometheus client library directly. The issue requests native OTel APIs with exporter compatibility (Prometheus, OTEL Collector, Grafana). This would be a refactor of metrics.py to use opentelemetry-api / opentelemetry-sdk.

    2. Feature Service name label on RED metrics — Currently RED metrics are broken out by feature_view but not by feature_service. Adding this with cardinality safeguards (configurable allowlist, default off) would complete the Feast-specific labeling story.

    3. Configurable feature count bins — The feature_count label on latency histograms currently uses raw counts. The issue proposes configurable bins (e.g., 1–10, 11–50, 51–200, 201+) to reduce cardinality while preserving useful segmentation.

    4. HTTP status class segmentation — Error metrics currently use status=success/error. Breaking this into HTTP status classes (2xx/4xx/5xx) would provide finer-grained error visibility.

    5. Drift detection — PSI, KS statistic, JS divergence, Wasserstein distance, mean shift (z-score). The DQM storage and computation infrastructure is in place; drift detection would build on top of the stored histograms and baselines. This is the most substantial remaining piece.

    @anshishrivastava — items 1–4 are good standalone PRs if you'd like to pick any up. Item 5 (drift detection) is a larger effort that builds on the DQM system in #6202. Happy to discuss approach on any of these.

  8. anshishrivastava commented on Jul 23, 2026

    @anshishrivastava
    Contributor

    @ntkathole @jyejare - Created issues for all pending items above

    1. Native OpenTelemetry SDK migration
    2. Configurable feature count bins
    3. Feature Service name label on RED metrics
    4. HTTP status class segmentation

    Please feel free to have a look at the issue details and proposed solutions and assign them to me and I'd be happy to take them up.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions