Repository navigation
Improve Feature Server Observability #5920
Description
Activity
Hi @franciscojavierarceo — I'd like to own this issue.
Before implementation, I'll prepare a lightweight design doc covering:
Metrics Design
- RED metrics per endpoint: request rate, error rate (by status code class), duration (p50/p95/p99)
- Labeling strategy:
feature_service,feature_view,entity_key_countwith cardinality controls to prevent label explosion in high-throughput production environments - Histogram buckets: logarithmic bucketing for latency (e.g., 1ms-10s range) tailored to typical online serving SLOs
- Error taxonomy: distinguish between client errors (invalid requests, missing features), server errors (registry/provider failures), and upstream errors (offline store timeouts)
Implementation Considerations
- OTel integration: native OpenTelemetry instrumentation with semantic conventions for observability, ensuring compatibility with Prometheus (via OTLP bridge), Datadog, and cloud-native collectors
- Performance impact: async metric recording to avoid adding latency to the critical serving path; batch aggregation where applicable
- Registry-aware labeling: dynamic label extraction from the feature store registry to maintain consistency with Feast's metadata model
Proposed Implementation Plan
- Middleware instrumentation + handler-level metric emission
- Unit tests (metric correctness) + integration tests (end-to-end validation)
- Example Grafana dashboard + basic alerting rules + quickstart docs
Happy to iterate on the design before implementation.
Please assign if this approach aligns with the roadmap. Thanks!Thank you @anshishrivastava for showing interest in it. Overall I see it aligns.
Feel free to tag @jyejare if you have some questions or need any clarification.Reacted by Anshi ShrivastavaHello @anshishrivastava , Thanks for showing your interest in the issue. Currently, we are in the process of developing RED metrics in the repo for OpenDataHub purpose; the PR will be up soon. But we are not covering all the metrics detailed in your design.
Sorry to say this but, I would want you to hold until implementation is done from our side, and then build on top it later.
Thanks for your understanding and patience. I am happy to answer any queries you have.
Hi @jyejare, checking in on the status of the OpenDataHub RED metrics PR.
Has that work been merged yet? I'm still very interested in building out the remaining metrics and Feast-specific labels we discussed (like
feature_serviceandfeature_viewsegmentation) to complete the observability story.If the base PR is up, I can start reviewing it or prepare a follow-up PR on top of it. Let me know how you'd like me to proceed!
@anshishrivastava Hey, thanks for the follow up. The work is postponed for a few days for other priorities, I shall keep you posted or at least I ll provide the top up work to you.
Reacted by Anshi Shrivastava@jyejare can you please confirm if this is available now? Please also add If something can be still worked on
@ntkathole @anshishrivastava — Here's an update on the current state.
What's available now
The core Feature Server observability is in place. The following have been merged:
Capability PR Status Online store RED metrics (request rate, error rate, latency histograms per endpoint) Pre-existing in metrics.py✅ Available Offline store RED metrics ( feast_offline_store_request_total,_latency_seconds,_row_count)#6340 ✅ Available SOX audit logging (online + offline, structured JSON via feast.auditlogger)#6340 ✅ Available Online store read/write duration histogram Pre-existing ✅ Available On-demand transformation duration histogram Pre-existing ✅ Available Materialization duration + result metrics Pre-existing ✅ Available Feature quality monitoring (DQM) — full backend, 8 offline stores, REST API, CLI, auto-baseline #6202 ✅ Available DQM UI (monitoring dashboard with histograms, time-series, filters) #6422 ✅ Available Batch + Log data source support for DQM #6202 ✅ Available What can still be worked on
The following items from this issue are open for contribution:
-
Native OpenTelemetry SDK migration — Current metrics use the Prometheus client library directly. The issue requests native OTel APIs with exporter compatibility (Prometheus, OTEL Collector, Grafana). This would be a refactor of
metrics.pyto useopentelemetry-api/opentelemetry-sdk. -
Feature Service name label on RED metrics — Currently RED metrics are broken out by
feature_viewbut not byfeature_service. Adding this with cardinality safeguards (configurable allowlist, default off) would complete the Feast-specific labeling story. -
Configurable feature count bins — The
feature_countlabel on latency histograms currently uses raw counts. The issue proposes configurable bins (e.g.,1–10,11–50,51–200,201+) to reduce cardinality while preserving useful segmentation. -
HTTP status class segmentation — Error metrics currently use
status=success/error. Breaking this into HTTP status classes (2xx/4xx/5xx) would provide finer-grained error visibility. -
Drift detection — PSI, KS statistic, JS divergence, Wasserstein distance, mean shift (z-score). The DQM storage and computation infrastructure is in place; drift detection would build on top of the stored histograms and baselines. This is the most substantial remaining piece.
@anshishrivastava — items 1–4 are good standalone PRs if you'd like to pick any up. Item 5 (drift detection) is a larger effort that builds on the DQM system in #6202. Happy to discuss approach on any of these.
Reacted by Anshi Shrivastava-
@ntkathole @jyejare - Created issues for all pending items above
- Native OpenTelemetry SDK migration
- Configurable feature count bins
- Feature Service name label on RED metrics
- HTTP status class segmentation
Please feel free to have a look at the issue details and proposed solutions and assign them to me and I'd be happy to take them up.
Is your feature request related to a problem? Please describe.
The Feature Server currently has limited built-in observability. While there is initial OpenTelemetry (OTEL) support, it does not expose standard service-level RED metrics (Rate, Errors, Duration) needed to operate the server in production.
Users need visibility into throughput, error rates, and latency for the core APIs:
/get-online-features/retrieve-online-documents/push/write-to-online-store/materialize/materialize-incrementalToday this often requires custom middleware or external tooling.
Describe the solution you'd like
Extend Feature Server’s OTEL integration to emit standard RED metrics out of the box, with useful Feast-specific breakdowns.
Metrics per endpoint should include:
Additional breakdowns:
1–10,11–50,51–200,201+).Implementation note (based on current
feature_server.pystructure):endpoint+status_class).request.statefor the middleware to include./materialize*), ensure duration histograms support multi-second/minute latencies.Optionally expose basic internal timings via tracing spans and/or additional histograms:
Metrics should use native OpenTelemetry APIs and work with common OTEL collectors/exporters (e.g., Prometheus, OpenTelemetry Collector, Grafana).
Describe alternatives you've considered
Additional context
Standard RED metrics with Feast-aware labels would significantly improve Feature Server operability and align Feast with common production observability practices, while keeping label cardinality under control.