DUE TO SPAM, SIGN-UP IS DISABLED. Goto Selfserve wiki signup and request an account.
...
Solr currently uses the Java Dropwizard 4 library for metric collection and event measurement. While Dropwizard is reliable, its primary limitation is the lack of support for tag-based or attribute-based metrics. This constraint reduces metric granularity because metrics are static, with tags embedded directly into metric names. To get around this Solr includes a Prometheus exporter that aggregates Dropwizard metrics into Prometheus using the /metrics api to scrape the metrics in prometheus format giving more dimensionality and easier aggregation with PromQL. However, this running an exporter introduces extra operational overhead and can be costly to maintain externally.
Furthermore, Solr faces vendor lock-in with Dropwizard. Transitioning toward an observability framework that inherently supports dynamic tags and attributes such as OpenTelemetry future-proof Solr against changing observability requirements.
Public Interfaces
The Open Telemetry Module or with the inclusion of SOLR-17432 Java Agent, takes the Open Telemetry API that is currently used for traces and embeds a OTLP exporter. can be used to export trace and now additional metrics. Users just need to set the following system property:
| Code Block |
|---|
otel.metrics.exporter=otlp |
With the Open Telemetry module or Java Agent, users can also connect with an OTel collector to export metrics to their chosen observability tool/vendor of their choice. An example of the OTel Collector config is shown below that will pull in OTLP and export into Prometheus on port `8889`.
| Code Block |
|---|
service:
# extensions: [health_check, pprof, zpages]
pipelines:
metrics:
receivers: [otlp]
exporters: [prometheus]
telemetry:
metrics:
address: 0.0.0.0:8888
level: detailed
receivers:
# Data sources: traces, metrics, logs
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
exporters:
prometheus:
endpoint: 0.0.0.0:8889
namespace: default |
Users can also user prometheus OTLP receiver as seen here https://prometheus.io/docs/guides/opentelemetry/ if they don’t want to use OTEL collector but this makes the prometheus server use a push mechanism.
The GET /admin/metrics API can change in a few ways. One way is outputting metrics in Prometheus format, the same as what `GET /admin/metrics?wt=prometheus` currently supports. This allows users to still use Prometheus servers or tools to pull and scrape solr nodes for metrics. This API will retain its filtering capabilities, allowing users to selectively expose the metrics they need. Other option is to bridge OTEL metrics into some form of JSON but will not necessarily look equivalent to dropwizard.
With this change, the Prometheus exporter will be deprecated unless repurposed for other uses. This will reduce maintenance of the project of no longer needing to support it. Users will be directed to alternative solutions such as the OTEL collector for transferring metrics to their preferred monitoring systems or vendors. For those requiring pre-aggregation filtering or custom solutions, tools like Telegraf can be recommended.
Potentially include OpenTelemetry Prometheus Exporter, enabling an embedded Prometheus server to bind to a separate port and expose Prometheus metrics for scraping. This can be set with the following system properties or environment variables in Solr:
| Code Block |
|---|
otel.metrics.exporter=prometheus
otel.exporter.prometheus.port=9464
otel.exporter.prometheus.host=0.0.0.0 |
Proposed Changes
There have been a number of problems with Solr still using Dropwizard and motivation to move off of it:
No tag/attribute metrics support for aggregation
Metrics with Dropwizard have no concept of tags as the tags themselves are embedded in the metrics. See the example below if someone wanted to look for the number of select requests from the /admin/metrics API:
| Code Block |
|---|
"QUERY./select.requests":0,
"QUERY./select.serverErrors":{
"count":0,
"meanRate":0.0,
"1minRate":0.0,
"5minRate":0.0,
"15minRate":0.0
} |
In order to retrieve this programmatically or just as a user, they need to rely on the format of the metrics which is in this example <category>.<handler>.<type>. Solr has filtering in a number of different ways such as regex to retrieve this with the following QUERY\./select\..*. If the metric name changes at all, this breaks that workflow.
In a tag based metric framework, appending or removing tags not relating to their metric query workflow would avoid this problem.
Complex and difficult filtering
Adding to the point above, filtering with Dropwizard is difficult and complex. The Prometheus Exporter shipped with Solr has a very large and complex JQ config file with a templating feature that is not easy to understand. For example, if a user wants all metrics relating to core in a core registry, here is the JQ query template from the prometheus exporter configuration:
| Code Block |
|---|
<template name="core" defaultType="COUNTER">
.metrics | to_entries | .[] | select(.key | startswith("solr.core.")) as $parent |
$parent.key | split(".") as $parent_key_items |
$parent_key_items | length as $parent_key_item_len |
(if $parent_key_item_len == 3 then $parent_key_items[2] else "" end) as $core |
(if $parent_key_item_len == 5 then $parent_key_items[2] else "" end) as $collection |
(if $parent_key_item_len == 5 then $parent_key_items[3] else "" end) as $shard |
(if $parent_key_item_len == 5 then $parent_key_items[4] else "" end) as $replica |
(if $parent_key_item_len == 5 then ($collection + "_" + $shard + "_" + $replica) else $core end) as $core |
$parent.value | to_entries | .[] | {KEYSELECTOR} as $object |
$object.key | split(".")[0] as $category |
$object.key | split(".")[1] | rtrimstr("]") | split("[") | .[0] as $handler | .[1] // "false" as $internal |
select($handler | startswith("/")) |
{METRIC} as $value |
if $parent_key_item_len == 3 then
{
name: "solr_metrics_core_{UNIQUE}",
type: "{TYPE}",
help: "See following URL: https://solr.apache.org/guide/solr/latest/deployment-guide/metrics-reporting.html",
label_names: ["category", "handler", "internal", "core"],
label_values: [$category, $handler, $internal, $core],
value: $value
}
else
{
name: "solr_metrics_core_{UNIQUE}",
type: "{TYPE}",
help: "See following URL: https://solr.apache.org/guide/solr/latest/deployment-guide/metrics-reporting.html",
label_names: ["category", "handler", "internal", "core", "collection", "shard", "replica"],
label_values: [$category, $handler, $internal, $core, $collection, $shard, $replica],
value: $value
}
end
</template>
|
The query is complex and brittle due to tags in its name making backwards compatibility difficult and easy to introduce regressions.
Maintenance and operational overhead of the prometheus exporter
The prometheus exporter is a Solr specific process for getting metrics from Dropwizard and requires constant maintenance with any new metrics or big changes that Solr introduces. Also running this external process can be costly to maintain.
Proposed changes
OpenTelemetry module
...
We aim to leverage the Solr Open Telemetry module or the OTel Java agent with the introduction of SOLR-17432, which initializes the Open Telemetry SDK for traces, to also push OTel metrics via OTLP. Metrics is currently hard coded to be off in the OTel module and we will remove this for metrics. We will collect metrics with the Open Telemetry API and this transition will completely remove Dropwizard, enabling support for an attribute-based metric system.
Open Telemetry supports different types of instruments which we will try to migrate from Dropwizards dropwizards equivalent. See Open Telemetry Meter section from its documentation.
...
The concept of registries goes away and metrics and meters will be retrieved from a GlobalOpenTelemetry object to measure events.
...
The GET /admin/metrics API will continue to exist, allowing users to scrape metrics. However, the removal of Dropwizard means the current format will change. We have several options:
- Replace API with Prometheus Formatted Metrics:
- The metrics API by default will expose Prometheus formatted metrics instead
- Leverage the Open Telemetry to prometheus bridge to go from OTEL -> Prometheus and use the Prometheus Response Writer.
- Deprecate all internal Prometheus Formatter code as we are no longer going from Dropwizard -> Prometheus
- Deprecate or repurpose the Prometheus Exporter.
- Create a Bridge/Shim for OTEL -> Dropwizard format:
- Export OTel metrics to Dropwizard JSON format to maintain backward compatibility with a custom bridge.
- Prometheus Exporter can continue to exist for custom configurations
Filtering
We will maintain filtering capabilities on the metrics API to enable users to limit the amount of data and metrics being scraped from the endpoint. Enhancements may include filtering by tags, preserving the concept of registries/groups with Open Telemetry, or specific metric names possibly. Need to see what the Open Telemetry SDK supports for filtering in Solr but also filtering is possible to happen at an exporter level with OTEL collectorlcollector/
Deprecation of Prometheus Exporter
With the removal and deprecation of the Prometheus Exporter, users will be encouraged to adopt alternative solutions such as the OTel collector. For those requiring pre-aggregation filtering or custom solutions, tools like Telegraf can be recommended. Deprecating this also is a benefit of no longer needing to maintain this. There are many open source solutions for exporters that I would argue does better than Solr's Prometheus Exporters and fits many other metrics pipelines. No need to re-invent the wheel.
JVM Metrics Collection
JVM metrics can be collected using the Open Telemetry runtime-telemetry-java17, which gathers comprehensive metric sets from JFR and JMX. We can also programmatically filter what metrics we want based JfrFeatures available giving further filtering for users. See the runtime-telemetry-java17 JfrFeature table.
...
| Code Block |
|---|
jvm_gc_duration_seconds_count{jvm_gc_action="end of minor GC",jvm_gc_name="G1 Young Generation",otel_scope_name="io.opentelemetry.runtime-telemetry-java17",otel_scope_version="2.14.0-alpha"} 3
jvm_gc_duration_seconds_sum{jvm_gc_action="end of minor GC",jvm_gc_name="G1 Young Generation",otel_scope_name="io.opentelemetry.runtime-telemetry-java17",otel_scope_version="2.14.0-alpha"} 0.013224165999999999
jvm_gc_duration_seconds_bucket{jvm_gc_action="end of minor GC",jvm_gc_name="G1 Young Generation",otel_scope_name="io.opentelemetry.runtime-telemetry-java8",otel_scope_version="2.14.0-alpha",le="0.01"} 3
jvm_gc_duration_seconds_bucket{jvm_gc_action="end of minor GC",jvm_gc_name="G1 Young Generation",otel_scope_name="io.opentelemetry.runtime-telemetry-java8",otel_scope_version="2.14.0-alpha",le="0.1"} 3
jvm_gc_duration_seconds_bucket{jvm_gc_action="end of minor GC",jvm_gc_name="G1 Young Generation",otel_scope_name="io.opentelemetry.runtime-telemetry-java8",otel_scope_version="2.14.0-alpha",le="1.0"} 3
jvm_gc_duration_seconds_bucket{jvm_gc_action="end of minor GC",jvm_gc_name="G1 Young Generation",otel_scope_name="io.opentelemetry.runtime-telemetry-java8",otel_scope_version="2.14.0-alpha",le="10.0"} 3
jvm_gc_duration_seconds_bucket{jvm_gc_action="end of minor GC",jvm_gc_name="G1 Young Generation",otel_scope_name="io.opentelemetry.runtime-telemetry-java8",otel_scope_version="2.14.0-alpha",le="+Inf"} 3
jvm_gc_duration_seconds_count{jvm_gc_action="end of minor GC",jvm_gc_name="G1 Young Generation",otel_scope_name="io.opentelemetry.runtime-telemetry-java8",otel_scope_version="2.14.0-alpha"} 3 |
Metric Reporters
Needs more research but possibly this goes away as well?
Grafana Dashboard
Consider introducing a new default Grafana dashboard to visualize metrics (potentially out of scope).
Compatibility, Deprecation, and Migration Plan
This change represents a significant update and will result in a break in backward compatibility, making it suitable only for major version releases. Users and applications relying on the metrics API will face compatibility issues, even if an OTel to Dropwizard shim is created, due to potential differences in naming conventions.
...
Use-case migration with Open Telemetry
Metrics API (GET /admin/metrics) and filtering
The GET /admin/metrics API will continue to exist, allowing users to scrape metrics with a pull based system. However, the removal of Dropwizard means the current format, its naming conventions and usage will change. The endpoint will now output Prometheus standard formatted metrics and filters will be around tags instead of regex. Below are some use-cases and how the user would migrate:
The endpoint is used an on demand curl to see Solr’s current state
For example, users who curl Solr with the following to see number of /select requests and errors:
| Code Block |
|---|
curl 'localhost:8983/solr/admin/metrics?regex=QUERY\./select\..*' |
With the new endpoint changes, the curl would look more like something below:
| Code Block |
|---|
curl 'localhost:8983/solr/admin/metrics?name=solr_requests_total&catergory=QUERY&handler=/select' |
Programmatic reading of the endpoint parsing JSON or XML
This is breaking and will no longer be supported. Users applications need to change and follow Prometheus Exposition format and data model seen here instead https://prometheus.io/docs/concepts/data_model/
Basic usage of the Prometheus Exporter with default configuration
Prometheus Exporter will be deprecated. If users are using the standard Solr configuration file for the prometheus exporter they can instead scrape directly from Solr nodes with a prometheus server or their application.
Custom usage of the Prometheus Exporter
Users who use a more custom configuration file for the prometheus exporter for complex prefilter or aggregation should instead adopt other third party tooling such as the OTel collector or Telegraf which also offer far more powerful aggregation and filter methods
An example of the OTel Collector config is shown below that will pull in OTLP and export into Prometheus on port `8889`.
| Code Block |
|---|
service:
# extensions: [health_check, pprof, zpages]
pipelines:
metrics:
receivers: [otlp]
exporters: [prometheus]
telemetry:
metrics:
address: 0.0.0.0:8888
level: detailed
receivers:
# Data sources: traces, metrics, logs
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
exporters:
prometheus:
endpoint: 0.0.0.0:8889
namespace: default |
Users who use a more custom configuration file for the prometheus exporter for complex pre-filtering or aggregation should instead adopt other third party tooling in the OTel collector or other tools such as Telegraf which also offer far more powerful methods and plugins.
Users can also user prometheus OTLP receiver as seen here https://prometheus.io/docs/guides/opentelemetry/ if they don’t want to use OTEL collector.
Usage of the /admin/metrics?wt=prometheus endpoint
Scrape and pull metric models will not change. Only changes are new metrics, naming conventions, API will not face disruptions, but they may need to update their PromQL queries for dashboards due to changes in metric names and tags.
Security considerations
This needs more research. Does OTLP support in-transit encryption such as TLS? Metrics could contain detailed information about the state of Solr. Maybe the metrics API should be bound to a separate port with auth and TLS.
Test Plan
Describe in few sentences how the SIP will be tested. We are mostly interested in system tests (since unit-tests are specific to implementation details). How will we know that the implementation works as expected? How will we know nothing broke?
[TBD]
Rejected Alternatives
Integration tests
Solr uses the dropwizard endpoint for some integration testing. For example, the PeerSyncReplicationTest asserts on REPLICATION.peerSync.errors for any failures on its integration testing. This workflow will stay the same, but developers should pull an Open Telemetry metric reader to find these metrics to assert on for their tests.
Metric reporters
Metric reporters will be deprecated. (TBD need more research and information on reporters)
Push model with OTLP
Users will not be open to the option of using Open Telemetry and its push model with OTLP metrics, similar to how traces and spans are being exported.
- Open Telemetry module of OTel Java agent
Users can enable the Open Telemetry module to push metrics or use industry standard plugins to work with their metric pipelines as a new optionIf there are alternative ways of accomplishing the same thing, what were they? The purpose of this section is to motivate why the design is the way it is and not some other way.