Visitar URL original
Releases · netdata/netdata · GitHub
Skip to content

Releases: netdata/netdata

v2.12.1

Choose a tag to compare

@netdatabot netdatabot released this 08 Oct 16:52

Netdata v2.12.1 is a patch release to address issues discovered since v2.12.0.

This release fixes a startup crash and a timezone detection bug in the Windows Agent, stabilizes the eBPF collector across a wider range of kernels and libbpf versions, reduces the load the Docker collector places on the Docker daemon, and removes unnecessary work from dbengine datafile deletion on busy Parents.

Windows

  • Fixed the Windows Agent crashing at startup with an access violation inside msys-2.0.dll. The bundled MSYS2 runtime carried an incomplete backport of an upstream Cygwin change; packages now ship a runtime rebuilt with the missing upstream fixes, the MSYS2 installer is pinned by tag and checksum, and packaging fails if the expected runtime is not the one shipped (#24170, @stelfrag)
  • Fixed Windows timezone detection reporting garbage, or reading past its buffer, when the Windows timezone had no match in the IANA mapping. The mapping is now parsed by attribute with bounded copies, the user's region is preferred over the world default as intended, and on failure the existing timezone is kept (#24101, @stelfrag)

eBPF collector

  • Synced the eBPF collector with ebpf-co-re v1.7.0.3 and kernel-collector v1.7.0.4, reworking the disk and process probes for better kernel compatibility, including hosts without BTF (#24008, @thiagoftsm)
  • Fixed ebpf.plugin aborting with heap corruption on hosts where the number of possible CPUs exceeds the number of online CPUs, such as VMware guests (#24008, @thiagoftsm)
  • Fixed a crash when an eBPF object or BTF data could not be loaded on builds with older libbpf versions (#24008, @thiagoftsm)
  • Fixed ebpf.plugin running the wrong action for most command-line options, so options such as --unittest now do what they say (#24008, @thiagoftsm)

Database

  • Removed a scan of the open cache that ran on every dbengine datafile deletion only to feed a debug check compiled out of release builds. The scan walked the whole clean queue under its lock, which can be large on busy Parents (#24162, @stelfrag)

Collectors

  • Docker: reduced the CPU load the collector causes in dockerd and containerd, which was significant on Docker Engine 29+ with the containerd image store. The image list is now refreshed at most once every 5 minutes, and the module's default update_every is raised from 1 to 10 seconds (#24146, @ilyam8)
  • Docker: fixed image usage metrics on daemons that negotiate an API older than 1.51, where every image was counted as active and dangling images were always reported as 0. Usage is now derived from the container list on every collection (#24147, @ilyam8)
  • Redfish: the configuration form now shows the username and password only for the authentication methods that use them (#24106, @ilyam8)
  • Redfish: fixed the collector permanently skipping resources whose @odata.id differs from the URI they were fetched from, such as the trailing-slash variants and remapped NetworkAdapters and Storage paths returned by HPE iLO. Affected jobs stayed in a partial state with a standing redfish_collection_status warning. Same-origin aliases are now accepted, the other URI safety checks still apply, and rejections report the specific reason (#24186, @ilyam8)
  • Ceph: removed a stray option selector from the configuration form, and stopped the Redfish TLS settings from rendering a duplicate tab strip (#24108, @ilyam8)
  • SMBIOS Memory: the collector is now built only on Linux, the only platform that exports the SMBIOS tables it reads. On macOS, FreeBSD and Windows its stock job used to start and fail with read SMBIOS entry point: ... no such file or directory (#24169, @ilyam8)

Support options

As we grow, we stay committed to providing the best support ever seen from an open-source solution. Should you encounter an issue with any of the changes made in this release or any feature in the Netdata Agent, feel free to contact us through one of the following channels:

  • Netdata Learn: Find documentation, guides, and reference material for monitoring and troubleshooting your systems with Netdata.
  • GitHub Issues: Make use of the Netdata repository to report bugs or open a new feature request.
  • GitHub Discussions: Join the conversation around the Netdata development process and be a part of it.
  • Community Forums: Visit the Community Forums and contribute to the collaborative knowledge base.
  • Discord Server: Jump into the Netdata Discord and hang out with like-minded sysadmins, DevOps, SREs, and other troubleshooters. More than 2000 engineers are already using it!

v2.12.0

Choose a tag to compare

@netdatabot netdatabot released this 30 Sep 15:24

NOTICE

The v2.12.0 build of the Netdata Agent for Windows has been discovered to have serious stability issues that cause it to crash randomly on startup due to an upstream bug in some of the tooling we depend on. As such, we have removed the installer for it from the release page. A fix is expected in v2.12.1, but until that release, users are encouraged to use the latest known working version (v2.11.1), which can be downloaded here.

Table of Contents

Release Summary

v2.11.0 brought network topology and made logs a first-class signal next to metrics. v2.12.0 adds OpenTelemetry
traces, out-of-band hardware monitoring through Redfish, and a diagnostics layer that makes SNMP topology debuggable,
alongside a large round of collector, query and stability fixes.

The third observability signal arrives. Netdata already ingested OpenTelemetry metrics and logs. This release adds
traces: OTLP trace records are received, normalized and stored by the agent under their own retention policy, with a
purpose-built columnar store — trace-aware indexes, bloom filters and rollups — and an otel-traces query Function that
searches traces, aggregates them over a window and drives an overview heatmap. You explore them in a new Traces tab
of the dashboard, available today as an Early Access feature.

Hardware monitoring goes out-of-band. The new Redfish collector talks to the BMC — iDRAC, iLO and any
DMTF Redfish implementation — rather than to the operating system, so as long as the BMC stays powered and reachable it
keeps reporting fans, temperatures, power supplies, drives and controllers when the host itself is unhealthy, rebooting or
powered down. It ships with sensor and hardware
inventory Functions, independent per-sensor threshold health, and an on-demand BMC log viewer. Alongside it,
SMBIOS Memory compares the firmware's per-slot memory inventory with an accepted baseline and raises an alert when a
recorded DIMM goes missing or comes back smaller — memory loss that a total-RAM figure alone does not pin to a slot.

The network work from v2.11.0 becomes operable. The SNMP topology engine gained an entire diagnostics layer: every
walk records its timing, its profile decisions and its failures; the evidence is retained in bounded form, published
independently of whether topology collection succeeded, packaged into a portable archive, and folded into the agent's
support bundle. A topology graph can then be replayed hermetically from that archive and inspected offline, which turns
"SNMP topology is not seeing this switch" from a site visit into a file you can send us. LLDP-V2 for PAN-OS, Meraki L2,
IpAddressTable support and a set of reusable topology-role profiles widen what the engine can see in the first place.

In Netdata Cloud, Netdata AI learns your infrastructure. With Infrastructure Knowledge, your team writes down
what only it knows — which services matter, what is expected, what normal looks like — and Netdata AI remembers facts
you tell it in conversation, then reads both before every conversation, investigation and report. AI conversations now
survive a closed tab and no longer fail when they grow long, and nodes being retired can be decommissioned so they
stop sending notifications.

And a lot of the release is consolidation. The Ceph collector was rebuilt as a complement to the Prometheus
collector and migrated to the v2 collector framework, with alert coverage and Tentacle support. go.d's secret handling
was systematically de-privileged: file and command secret providers now run without elevated privileges and through
nd-run, and secret references in auto-discovered jobs are no longer resolved unless the discovery pipeline is
explicitly trusted. Dynamic Configuration now reports an accepted go.d change separately from whether the job,
discovery pipeline or secret store has started, so clients should check configuration status for the outcome. Six independent correctness fixes landed in the query engine, and a long run of dbengine, SQLite,
dictionary, ACLK and ML fixes closes out crashes and leaks — many of them already shipped to users in the 2.11.1 patch
release and consolidated here.

Release Highlights

OpenTelemetry Traces, in Early Access

The agent's OTLP receiver now accepts traces in addition to metrics and logs, and a new Traces tab in the
dashboard lets you explore them.

Trace records arriving on the OTLP/gRPC endpoint are normalized, written to the write-ahead log and sealed
into the same columnar store that serves logs, extended for spans: span-extras columns, a trace-aware row index, trace
bloom filters, span combination and trace rollups. Queries are served by a new otel-traces Function, which exposes
capability discovery plus a trace search that filters on span fields and breaks each trace's spans down by service, a
window aggregate embedded in the search response, and an overview heatmap that follows the page's current selections.

Seeing your traces. Pick a time window and the Traces tab shows you where the time goes: a heatmap of trace
durations, the traces themselves with the services each one touched and any errors along the way, and filters to narrow
things down by service, span name or minimum duration. Click a trace to see its spans laid out as a waterfall, then
click a span for its attributes, resource, scope, events and links. You'll find Traces under Explore in a room, and
right next to Logs on a node's page.

Note

Traces is still in Early Access, so it stays hidden until you switch it on: click Early Access in the sidebar
and flip the Traces toggle (your browser remembers the choice). Just like logs, you'll need to be signed in to
Netdata Cloud, with access to the node's Space, to see them.

Traces are stored under their own retention settings — the traces section of otel.yaml takes the same rotation and
retention options as logs. With remote storage enabled, trace files that local retention has already removed are read
back from object storage for queries, as log files are, through a download cache that both signals now share at
{base_dir}/remote-read. A remote file that cannot be downloaded marks the answer as partial (remote_unavailable),
while a query that needs more remote data than the cache can hold fails with a request to narrow the range or raise
remote_storage.read_cache_max_size. Offloaded files are downloaded whole and one at a time, so the first query over
offloaded history can be slow, and a lookup of one trace by ID without a time range searches everything kept, local and
offloaded, so it fails once the offloaded history is larger than the cache.

Related work hardened the plugin around it: the otel-plugin now binds its gRPC endpoint before declaring its
Functions. On a port conflict it exits once without advertising a receiver it cannot listen on, and the agent disables
it until the next agent restart instead of restarting it in a loop. The OpenTelemetry documentation gained a
known-errors section covering exactly that case.

(#23479, #23738,
#23739, #23826,
#23905, #24088)

Redfish: Out-of-Band Server Hardware Monitoring

The new go.d/redfish collector monitors server hardware through the DMTF Redfish API of a baseboard management
controller. Because it queries the BMC rather than the operating system, it keeps reporting while the host is
unresponsive, mid-reboot or powered off, as long as the BMC itself stays powered and reachable — which is precisely when
hardware questions get asked.

One job per BMC covers the chassis and system inventory: thermal sensors and fans, power supplies and consumption,
processors and memory, storage controllers and drives, and the overall health roll-up.

Beyond charts, it ships operator tooling:

  • Sensor and hardware Functions — on-demand tables of the full sensor set and the hardwa...
Read more

v2.11.1

Choose a tag to compare

@netdatabot netdatabot released this 16 Sep 16:35

Release notes

Netdata v2.11.1 is a patch release to address issues discovered since v2.11.0.

This is a larger-than-usual patch release focused on query correctness, database engine stability, Cloud connectivity robustness, health/alert efficiency, and shutdown and process-lifetime safety, along with a number of collector and platform fixes, including broader SQL Server coverage and Windows CPU and disk improvements.

Query engine

  • Preserved exact SUM values when a stored point spans result rows or an automatic storage-tier boundary, and preserved incremental-sum baselines so buckets holding only an opening sample no longer produce empty queries (#23512, #23505, @ktsaou)
  • Kept API timestamps independent of each Agent's retention, so the same explicit or relative request returns the same row timestamps across Agents (#23507, @ktsaou)
  • Attributed tier-0 RESET and anomaly evidence to the result row that actually contains the source sample, and fixed anomaly-rate contributor counts in two-pass group-by queries (#23504, #23508, @ktsaou)
  • Applied options=absolute to points fetched from the incoming tier during a plan switch, so a negative value can no longer cross a tier transition unchanged (#23503, @ktsaou)
  • Fixed a logical operator in the JSON wrapper so unqueried dimensions are excluded as expected unless all dimensions are requested (#23639, @stelfrag)

Database engine

  • Fixed a use-after-free and a fatal in merged extent queries by acquiring page-cache references before publishing page details and synchronizing merged extent query teardown under the extent spinlock (#23475, #23521, @stelfrag)
  • Fixed a double free of stale MRG metrics during journal v2 migration, and an empty flush-batch counter leak that could leave shutdown or datafile rotation waiting indefinitely (#23560, #23699, @stelfrag)
  • Prevented negative and epoch-based tier retention reporting: no 32-bit truncation when formatting durations, unknown or zero start times treated as unavailable, empty journals ignored, and extrapolated retention capped (#23648, @stelfrag)
  • Accepted valid page cadences longer than one day instead of treating them as invalid and emitting false repair telemetry (#23514, @ktsaou)
  • Avoided UUID map create/free churn and partition write-lock contention during MRG metric lookups, and included full cache and page identity in PGC fatal messages to make cache issues diagnosable (#23538, #23522, @stelfrag)
  • Updated the bundled SQLite to 3.53.4 (#23511, @stelfrag)
  • Repaired the -W createdataset and dbengine stress-test workloads, which crashed on a NULL labels allocator, stamped every generated sample with the wall clock instead of the requested time range, and double-freed state at host teardown (#23758, @stelfrag)

Cloud connectivity (ACLK)

  • Added TLS hostname and IP verification for ACLK HTTPS requests, while preserving insecure-mode behavior (#23463, @stelfrag)
  • Stopped reporting waived self-signed certificates as certificate verification failures in insecure mode, so later I/O or HTTP errors are reported for what they are (#23732, @stelfrag)
  • Switched the HTTPS and MQTT clients to monotonic, microsecond-based timeouts with an I/O progress watchdog, and bounded pre-CONNACK polling to one second so a quiet peer can no longer delay shutdown by up to 60 seconds (#23105, #23733, @stelfrag)
  • Stopped counting missing MQTT PUBACKs as successful acknowledgements, tracking them with a dedicated timeout counter that is now preserved across disconnects without touching a freed client (#23486, #23734, @stelfrag)
  • Preserved pooled response buffers and their metadata for compressed ACLK responses, and fixed WebSocket masking of payloads spanning ring-buffer boundaries (#23485, @stelfrag)

Health and alerts

  • Fixed the Windows 10min_cpu_usage alert attaching to service.*_cpu_utilization charts and firing false CRITICAL alarms (#23559, @thiagoftsm)
  • Indexed alerts by name per host, replacing full-host alert scans during health-variable resolution (#23554, @stelfrag)
  • Avoided re-initializing alert prototypes when chart metadata has not changed, so continuous child-chart re-registrations no longer trigger repeated health re-evaluation (#23515, @stelfrag)
  • Built raised-alert summaries only when a notification is actually sent, avoiding per-iteration work and a potential lock-order deadlock with chart removal (#23544, @stelfrag)
  • Fixed an incorrect error check on SQL statement preparation in the health database code, and removed two unused health_log_detail indexes to reduce parent database storage and B-tree maintenance (#23545, #23722, @stelfrag)

Metadata, labels and internals

  • Deferred dictionary destruction while dictionary APIs are in flight, making set/get/delete and traversal safe against concurrent destruction during shutdown (#23634, @stelfrag)
  • Fixed a use-after-free in dictionary garbage collection when a delete callback re-entered the collector and freed the cached successor mid-walk, and stopped view collections leaving master-deleted items behind (#23819, @stelfrag)
  • Stopped the obsolete-chart reaper from freeing a chart another subsystem still references, which left an unindexed chart holding its destroy lock and could fatal a collector still writing to it (#23761, @stelfrag)
  • Deferred machine-learning worker queue teardown until after all collectors have stopped, so a collector that outlives the shutdown deadline can no longer abort on a destroyed queue mutex (#23757, @stelfrag)
  • Reset and unblocked SIGPIPE in spawned children, so a child whose pipes are closed exits instead of surviving on EPIPE — which had stranded hundreds of powermetrics processes on macOS — reported failed signal delivery instead of dropping it silently, and made the spawn-server regression tests deterministic (#23751, #23815, @stelfrag)
  • Probed for KSM support before enabling page deduplication instead of assuming it, so unsupported kernels no longer report KSM errors, and limited KSM marking to the dbengine page pools that can actually merge (#23822, @stelfrag)
  • Skipped sanitizing and re-interning unchanged chart metadata on repeated chart registration, and incremented the label version only on actual label mutations (#23603, #23513, @stelfrag)
  • Fixed a text-buffer cleanup issue and a dyncfg tree sizing race during concurrent configuration updates (#23657, @stelfrag)
  • Replaced sprintf() with bounded snprintfz() in inicfg numeric formatting, and replaced dynamic /proc/net/dev path format strings with literal ones (#23638, #23450, @stelfrag)
  • Reduced pulse overhead by running mallinfo2() collection on a configurable, pulse-aligned interval, stamping it after completion so a slow call cannot monopolize the duty cycle, and stopped redundant pulse chart metadata version bumps ([#23579](https://github.com/netdata/netdata/pull...
Read more

v2.11.0

Choose a tag to compare

@netdatabot netdatabot released this 12 Aug 18:38

Table of Contents

Release Summary

Netdata has always been exceptionally good at one particular thing: telling you, per second and with no configuration,
what is happening inside a machine. v2.11.0 is the release where that stops being the boundary.

Two things happen here, and together they change what Netdata is.

First, Netdata learns the network. A new Network Monitor dashboard brings device inventory, metrics, topology,
flows, traps and alerts into one place. Behind it, Network Flows receives NetFlow, IPFIX and sFlow from your routers,
switches and firewalls and turns raw records into a faceted view of who is talking to whom. Network Topology walks
your devices over SNMP and draws the Layer 2 and Layer 3 map — LLDP neighbours, forwarding databases, OSPF and BGP
adjacencies — while Application Dependency Mapping does the same for your software, reading your Linux hosts'
kernels to map what each process, container and Kubernetes workload talks to, with nothing to instrument. The SNMP Trap
Listener
makes Netdata a trap receiver that decodes traps from 800+ vendor profiles into
charts and readable messages instead of raw OIDs. Feeding all of it, SNMP device coverage grew from 229 to 273
profiles
. The usual Netdata rules still apply: the profiles are already written, collection and storage stay on your own
infrastructure, and there is little to configure beyond telling Netdata where the devices are. Network Monitor, Network
Flows and Network Topology ship as technical previews — complete enough to run today, open to every Netdata user
while in preview, and still moving fast.

Second, and just as important, logs become a first-class pillar of the platform. This release adds OpenTelemetry
logs
— OTLP log records are ingested, stored and queried from the same Logs tab that already serves systemd journal
and Windows events, with journald, filelog and syslog receivers, configurable retention, log-to-metric conversion
and parsing for unstructured lines. The deeper change is architectural: the same faceted, high-cardinality log engine now
backs systemd journal, Windows events, OpenTelemetry logs, SNMP traps and network flows alike. Five very different
signal types, one query surface, all stored on your own infrastructure.

Around those two, this release completes the major-cloud set with a new AWS CloudWatch collector — the counterpart to
the Azure Monitor collector added in v2.10.0 — and rebuilds the Prometheus collector around a real relabeling engine
with application profiles. Underneath all of it sits the largest stability and memory-safety effort we have ever shipped
in a single release: 87 incremental parts across the database engine, streaming, ACLK, health, ML and libnetdata.

Netdata's strong endpoint monitoring capabilities have been enhanced. macOS monitoring was rebuilt natively: unified
logs through Apple's OSLog framework, GPU, SMC and IOHID sensors, fans, power and battery, thermal pressure and NVMe
health, none of which was reachable before without dropping to log show and powermetrics by hand. Windows added
Active Directory and SMB monitoring on a reworked perflib layer, FreeBSD added system-call monitoring, and
Linux added audit subsystem monitoring. One agent, four endpoint operating systems, per-second resolution on each.

In Netdata Cloud, AI troubleshooting just became more accurate and more relevant: you can now connect your own MCP servers
to your Netdata Space — GitHub, PagerDuty and Atlassian, and your own custom servers — and Netdata AI will use them while
investigating. Alongside that, new Fleet Management views render nodes as status hexagons, the Nodes and Alerts views were
rebuilt to stay fluent in 50,000-node rooms, silencing rules gained timezone-anchored scheduling, and the mobile app
opened up to everyone on free plans.

Feature Highlights Details
Network Monitor Dashboard One place for the whole network
Technical Preview
• Device inventory, vendors, health and top talkers at a glance
• Overview, Devices, Metrics, Topology, NetFlow, Traps and Alerts sub-tabs
• Per-device and per-interface traffic, speed and error rates
Network Flows NetFlow v5/v9, IPFIX, sFlow
Technical Preview
• New high-throughput flow-analysis plugin with its own four-tier journal
• Cisco ASA NSEL accounting
• Enrichment with cloud provider IP ranges, BGP/RIPE RIS data, and private-IP labelling
Network Topology L2 and L3 topology discovery
Technical Preview
• LLDP, FDB and STP based Layer 2 topology
• L3 subnet segments plus OSPF and BGP adjacencies
• MAC OUI vendor lookup and reverse DNS enrichment
Application Dependency Mapping Per-host service maps, no instrumentation
Technical Preview
• What each process talks to, read from the kernel socket table
• Attributed per container, image, systemd unit and Kubernetes pod/workload (Linux)
• Regroup the map by process name, container or PID
SNMP Trap Listener 800+ vendor trap profiles • Profile-based trap decoding into charts and log entries
• Trap enrichment and forwarding to SIEM
• Configurable via the Dynamic Configuration UI
Expanded SNMP Device Coverage 273 device profiles, up from 229 • Cisco Catalyst, Cisco Nexus, FortiGate and Check Point improvements
• BGP session monitoring and network-device licence monitoring
• SNMPv3 context names, bitmask value mappings, ping_only mode
Logs One engine, five signal types • OpenTelemetry (OTLP) log ingestion, query and retention
• journald, filelog and syslog receivers; log-to-metric conversion
• Same faceted engine now backs journal, Windows events, traps and flows
AWS CloudWatch Collector 47 service profiles, 800+ metrics • Automatic resource discovery with tag filtering and multi-account support
• Stock health alerts, incl. load-balancer target health and MSK
• Query policies to control API cost on sparse metrics
New Collectors Cato Networks, PAN-OS, and more • SASE and firewall monitoring for Cato Networks ...
Read more

v2.10.4

Choose a tag to compare

@netdatabot netdatabot released this 15 Jul 14:49

Release notes

Netdata v2.10.4 is a patch release to address issues discovered since v2.10.3.

This is a larger-than-usual patch release focused on stability, memory safety, and crash resilience across the database engine, streaming, Cloud connectivity, and collectors. It also backports the new macOS hardware sensor collectors to the 2.10.x line.

New in this release

  • macOS hardware sensor collectors: GPU power, clock, and temperatures, SMC and IOHID sensors and fans, power sources and thermal pressure, NVMe SMART, with a powermetrics fallback via ndsudo, plus a cross-OS sensors function and a shared temperature-histogram context (#22475, #23085, @ktsaou)
  • Bounded-cardinality process grouping for apps.plugin on macOS, keeping per-process monitoring meaningful on hosts with high process churn (#23085, @ktsaou)

Database engine and metadata

  • Hardened DBENGINE against corrupted or truncated data files: extent disk-size and uncompressed-page-bounds validation, SIGBUS protection and header-offset bounds checks on v2 journal walks, improved journal file access error handling, and spinlock-holder identification for datafile/journal deadlock diagnostics (#22324, #22514, #22310, #22725, @stelfrag; #22666, @jmestwa-coder)
  • Fixed page-cache races in pgc_page_add and pgc_queue_del, a deadlock on dimension creation, and a crash on virtual node takeover (#22466, #22400, #22417, @stelfrag)
  • Validated data loaded from SQLite to prevent crashes on corrupted databases, and improved UUID handling and error reporting in SQLite functions (#22679, #22233, @stelfrag)
  • Kept chart indexes allocated while a host is archived and guarded against a null root index in chart lookups, avoiding null dereferences during queries (#23074, #23056, @stelfrag)
  • Fixed an rrdcontext metadata leak on non-dbengine hosts (#22438, @stelfrag)
  • Rotated the active datafile and requeued the pending extent on unrecoverable write errors, so a failing disk no longer stalls the database engine (#23048, @ktsaou)

Streaming and replication

  • Tracked and accounted memory allocation size in replication queries, and fixed a sender replication counter leak on obsolete charts (#22756, #22428, @stelfrag)
  • Rejected oversized ZSTD frames and handled decompression errors gracefully (log and fail the connection instead of fatal()) (#22830, @stelfrag)
  • Fixed the streaming receiver discarding already-delivered data when a child disconnects (#23118, @ktsaou)

Cloud connectivity (ACLK)

  • Prevented rare unbounded one-core CPU spins in the cloud-connection loops, and avoided returning an uninitialized packet_id on publish failure (#22879, @ktsaou; #22504, @stelfrag)

Alerts and machine learning

  • Bounded the alert notification execution wait so a stuck notification script cannot block progress, and fixed a shutdown race when restoring alert information from the database (#22626, #22448, @stelfrag)
  • Added safeguards against ML database corruption and streamlined the recovery process (#22478, @stelfrag)

Collectors

  • Fixed systemd-journal.plugin memory retention after queries and its apps.plugin accounting (#23089, @ktsaou)
  • Fixed file descriptor accounting in apps.plugin (adding nfds monitoring) and eBPF FD PID map iteration (#22447, @arch-yunus; #22436, @stelfrag)
  • Fixed buffer overflows and counter-size issues in platform collectors: a stack buffer overflow in macOS mach_smi, FreeBSD counter size mismatches and an off-by-one in freebsd_ipfw and the claim code, and a freeipmi.plugin watchdog underflow at low system uptime (#22553, @artem; #23044, @DavidMarec; #22710, @vkalintiris; #22490, @ktsaou)
  • Fixed Windows Events row rendering crashes with stricter size checks, safer XML parsing, and robust variant handling (#22872, @ktsaou)
  • Restored stable temperature and power collection in go.d/nvidia_smi with NVIDIA driver 580 XML output variants (#23047, @copilot-swe-agent)
  • Fixed Windows hardware detection so virtual machines are no longer reported as bare metal (#22942, @thiagoftsm)
  • Moved the fail2ban socket path into ndsudo for go.d/fail2ban (#22745, @ilyam8)
  • Fixed a pluginsd cleanup race with an active collector, added a slot bounds check to the pluginsd parser, and initialized mountpoint state before the slow worker in the diskspace plugin (#22207, #22598, #22298, @stelfrag)

Security and API

  • The /api/v3/settings endpoint is now gated behind HTTP_ACL_DASHBOARD instead of HTTP_ACL_NOCHECK, enforcing connection allowlists and blocking anonymous state-changing requests (#22896, @stelfrag)
  • Fixed the csvjsonarray output format emitting invalid JSON when label-quotes is passed (#23115, @ktsaou)

Core stability and memory safety

Read more

v2.10.3

Choose a tag to compare

@netdatabot netdatabot released this 27 Apr 17:04

Release notes

Netdata v2.10.3 is a patch release to address issues discovered since v2.10.2.

This patch release provides the following bug fixes and updates:

  • Fixed a per-PID shared-memory pool leak in ebpf.plugin that, on hosts with normal process churn, filled the 32,768-slot pool within ~15 hours and then pegged a CPU core at 100% in an infinite map-iteration loop; aggregation paths now use a non-allocating lookup, freshly-allocated slots are zeroed, module bits are swept on exit, and stale shared-memory and semaphore objects are properly cleaned up on init (#22232, @ktsaou)
  • Switched the SNMP collector's primary uptime source to SNMP-FRAMEWORK-MIB::snmpEngineTime (in seconds) to avoid the ~497-day TimeTicks wrap of hrSystemUptime/sysUpTime, while keeping the existing systemUptime metric name and HR-MIB fallback (#22231, @ilyam8)
  • Split dynamic configuration job-name validation per domain so service discovery, vnode, and secret-store names can include dots (e.g., FQDNs), while collectors retain strict naming rules (#22247, @ilyam8)
  • Removed the unused extra_details field from the go.d/powerstore Hardware struct to fix /hardware response decoding errors that caused job check failures (#22291, @ilyam8)

Support options

As we grow, we stay committed to providing the best support ever seen from an open-source solution. Should you encounter an issue with any of the changes made in this release or any feature in the Netdata Agent, feel free to contact us through one of the following channels:

  • Netdata Learn: Find documentation, guides, and reference material for monitoring and troubleshooting your systems with Netdata.
  • GitHub Issues: Make use of the Netdata repository to report bugs or open a new feature request.
  • GitHub Discussions: Join the conversation around the Netdata development process and be a part of it.
  • Community Forums: Visit the Community Forums and contribute to the collaborative knowledge base.
  • Discord Server: Jump into the Netdata Discord and hang out with like-minded sysadmins, DevOps, SREs, and other troubleshooters. More than 2000 engineers are already using it!

v2.10.2

Choose a tag to compare

@netdatabot netdatabot released this 14 Apr 14:38

Release notes

Netdata v2.10.2 is a patch release to address issues discovered since v2.10.1.

This patch release provides the following bug fixes and updates:

  • Fixed ZFS-related crashes in diskspace.plugin by guarding a NULL filesystem field and replacing the blocking pool-capacity collector with a lightweight cache fed by existing statvfs calls, preventing coredumps on degraded or exporting ZFS pools (#22188, @thiagoftsm)
  • Reduced default SNMP MaxOIDs from 60 to 20 for more reliable polling, and removed 32-bit counter fallbacks from the IF-MIB profile to prevent counter type switching and overflow between collection cycles (#22203, @ilyam8)
  • Eliminated false-positive "timed out waiting for enable/disable decision" warnings by removing the wait-decision timeout and making the dyncfg command path non-droppable, so back-pressure flows upstream instead of producing 503 errors (#22201, @ilyam8)

Support options

As we grow, we stay committed to providing the best support ever seen from an open-source solution. Should you encounter an issue with any of the changes made in this release or any feature in the Netdata Agent, feel free to contact us through one of the following channels:

  • Netdata Learn: Find documentation, guides, and reference material for monitoring and troubleshooting your systems with Netdata.
  • GitHub Issues: Make use of the Netdata repository to report bugs or open a new feature request.
  • GitHub Discussions: Join the conversation around the Netdata development process and be a part of it.
  • Community Forums: Visit the Community Forums and contribute to the collaborative knowledge base.
  • Discord Server: Jump into the Netdata Discord and hang out with like-minded sysadmins, DevOps, SREs, and other troubleshooters. More than 2000 engineers are already using it!

v2.10.1

Choose a tag to compare

@netdatabot netdatabot released this 10 Apr 16:01

Release notes

Netdata v2.10.1 is a patch release to address issues discovered since v2.10.0.

This patch release provides the following bug fixes and updates:

  • Added ping_only option to the SNMP collector, allowing collection of only ICMP round-trip time metrics while still using SNMP at startup for device identification and labeling (#22180, @ilyam8)
  • Fixed SNMP collector retries serialization so a zero value is correctly preserved, and disabled bulk walk detection when MaxRepetitions is set to 0 (#22179, @ilyam8)
  • Reduced transient 503 errors on dynamic configuration enable/disable by buffering the command channel and extending the handoff timeout in the go.d plugin (#22183, @ilyam8)

Support options

As we grow, we stay committed to providing the best support ever seen from an open-source solution. Should you encounter an issue with any of the changes made in this release or any feature in the Netdata Agent, feel free to contact us through one of the following channels:

  • Netdata Learn: Find documentation, guides, and reference material for monitoring and troubleshooting your systems with Netdata.
  • GitHub Issues: Make use of the Netdata repository to report bugs or open a new feature request.
  • GitHub Discussions: Join the conversation around the Netdata development process and be a part of it.
  • Community Forums: Visit the Community Forums and contribute to the collaborative knowledge base.
  • Discord Server: Jump into the Netdata Discord and hang out with like-minded sysadmins, DevOps, SREs, and other troubleshooters. More than 2000 engineers are already using it!

v2.10.0

Choose a tag to compare

@netdatabot netdatabot released this 09 Apr 14:40

Table of Contents

Release Summary

Netdata v2.10.0 introduces secrets management, a Nagios plugins collector, Azure Monitor support, and broad stability improvements.

Feature Highlights Details
Secrets Management 4 resolver types, 4 backends • Environment variables, files, commands, and secretstore references
• AWS Secrets Manager, Azure Key Vault, GCP Secret Manager, HashiCorp Vault
Enhanced AI Conversations Reports and Alerts integration • Resume and expand generated AI reports
• Ask AI about active alerts or charts, their root causes, and recommended remediation steps
Advanced Alerting Silencing, Evaluation & Acknowledge • Complete support for recurrence rules for Alerts Silencing
• Test alert definitions against historical data before deploying
• Acknowledge alerts you're working on
Improved Custom Dashboards Complete Customisation • Add Metrics, Logs, Events, Node List, Alerts, Live Functions and Text components
• Duplicate Custom Dashboards
Nagios Plugins Collector Run any Nagios-compatible check • Automatic performance data charts
• Threshold-based alerting with soft/hard state logic
• Execution metrics (duration, CPU, memory)
Azure Monitor Collector 38 service profiles, 1,300+ metrics • Automatic resource discovery via Azure Resource Graph
• Multi-subscription, flexible scoping by resource groups/regions/tags
Azure AD for Database Collectors MSSQL, PostgreSQL, Generic SQL • Service principal, managed identity, and default credential chain
• Passwordless auth for Azure-hosted databases
New Storage Collectors Dell PowerStore, Dell PowerVault • Hardware health, performance, capacity, and sensor monitoring
• Built on V2 collector framework
Expanded Collector Coverage vSphere, MSSQL, SNMP, Docker • vSphere datastores/clusters/resource pools
• MSSQL Always On AG monitoring
• SNMP IPSec/VPN profiles for FortiGate, Juniper, MikroTik, Check Point
• Docker container listing function
OpenTelemetry Improvements Metrics pipeline overhaul • Proper slot-based aggregation with configurable intervals
• Multi-slot ingestion with out-of-order support
Stability & Performance Production reliability • Faster agent startup
• ML prediction optimization
• Alerts API speedup
• Multiple crash and race condition fixes

Release Highlights

Secrets Management

Keep collector credentials out of plain-text configuration files. Netdata now lets you reference secrets in collector configurations instead of storing them directly. Passwords, tokens, and API keys are resolved at runtim...

Read more

v2.9.0

Choose a tag to compare

@netdatabot netdatabot released this 16 Feb 16:37

Table of Contents

Release Summary

Netdata v2.9.0 brings powerful database observability and expanded OpenTelemetry support.

Feature Highlights Details
Interactive Query Analysis 14+ databases supported • Identify slow queries
• Debug bottlenecks
• Monitor operations
• No manual database connections required
Rebuilt SQL Server Collector Rewritten in Go • Comprehensive metrics
• Query Store integration
• UI-based configuration
OpenTelemetry Log Ingestion Logs via otel plugin • Stored in systemd-compatible journal files
• Configurable retention policies
Stability Improvements Production reliability • Under-the-hood fixes for more robust operation

Release Highlights

New Top Tab Functions: 14+ Databases and SNMP

We've added interactive query analysis functions to database collectors, letting you identify slow queries, long-running operations, and performance bottlenecks directly from the Netdata
dashboard—no need to connect to the database and run diagnostic queries manually.

Database Functions Description
ClickHouse Top Queries • Aggregated query stats from system.query_log
• Execution time, memory usage, rows read/written
CockroachDB Top Queries • Statement statistics from crdb_internal
• Execution counts, latency, rows processed
Running Queries • Currently executing statements via SHOW CLUSTER STATEMENTS
• Client info, duration, distributed execution status
Couchbase Top Queries • Completed N1QL requests from system:completed_requests
• Service time, result count, error tracking
Elasticsearch/OpenSearch Top Queries • Active search tasks from Tasks API
• Running time, search type, node distribution
MongoDB Top Queries • Slow operations from system.profile
• Execution time, docs examined, keys examined, plan summary
MS SQL Server Top Queries • Query Store statistics
• CPU time, logical reads/writes, memory grants, parallelism
Deadlock Info • Latest deadlock from system_health Extended Events
• Victim process, lock mode, wait resource
Error Info • Recent SQL errors from Extended Events session
• Error number, message, query text
MySQL/MariaDB Top Queries • Digest statistics from performance_schema
• Execution time, lock time, rows examined/sent
Deadlock Info • Latest InnoDB deadlock from SHOW ENGINE INNODB STATUS
• Victim transaction, lock mode, wait resource
Error Info • Recent SQL errors from Performance Schema history
• Error number, SQLSTATE, message per query digest
Oracle DB Top Queries • SQL statistics from V$SQLSTATS
• CPU time, elapsed time, buffer gets, disk reads
Running Queries • Active sessions from V$SESSION
• Wait events, blocking sessions, SQL text
PostgreSQL Top Queries • Statement statistics from pg_stat_statements
• Total/mean time, shared blocks hit/read, temp blocks
Running Queries • Active queries from pg_stat_activity
• Duration, wait events, client info, backend state
ProxySQL Top Queries • Query digest from stats_mysql_query_digest
• Execution time, rows affected/sent, errors
Redis Top Queries • Slow commands from SLOWLOG
• Command name, duration, client info
RethinkDB Running Queries • Active jobs from rethinkdb.jobs
• Query text, duration, involved servers
YugabyteDB Top Queries • YSQL statistics from pg_stat_statements
• Execution time, calls, rows processed
Running Queries • Active backends from pg_stat_activity
• Query state, wait events, elapsed time
Generic SQL User-Defined • Custom SQL functions defined in job configuration
• Interactive table views in Top tab for any SQL database

What You Can Do:

  • Identify resource hogs: Sort by total execution time, CPU time, or I/O
  • Spot frequent queries: Find high-call-count queries that may benefit from caching
  • Debug in real-time: See currently running queries and which client initiated them
  • No extra tools needed: All analysis happens in the Netdata dashboard's Top tab

Important

Some functions require database-specific configuration (e.g., enabling Query Store, pg_stat_statements, or profiling).

See each collector's documentation for prerequisites.

Also Added: SNMP Network Interfaces Function

Collector Function Description
SNMP Network Interfaces • Real-time interfa...
Read more