[New Package] Add Grafana integration - #17433
Conversation
New integration for collecting logs and Prometheus metrics from Grafana instances (10.x, 11.x, 12.x). Includes ingest pipelines for both JSON and logfmt log formats, field mappings that rename prometheus.metrics.* to a structured grafana.* namespace, two Kibana dashboards (overview and logs), and pipeline tests with data captured from a real Grafana 12.3.2 instance.
ReviewersBuildkite won't run for external contributors automatically; you need to add a comment:
NOTE: https://github.com/elastic/integrations/blob/main/.buildkite/pull-requests.json contains all those details. |
|
/test |
|
@claude for detailed initial review, as a new package all requirements to be met. |
|
Claude finished @lalit-satapathy's task in 3m 6s —— View job Detailed Review of New Grafana Integration PackageI've completed a thorough review of this new Grafana integration package. The review has been submitted formally on the PR. Summary:
See the formal review for complete details. Why is Claude responding? | Type Give us feedback! React with 🚀 if perfect, 👍 if helpful, 👎 if not. |
There was a problem hiding this comment.
This is a well-structured new integration package for Grafana. I've reviewed the package against integration requirements and found it to be comprehensive.
✅ Requirements Met
- Package manifest: Format version 3.3.0, proper categories (observability, monitoring), Kibana version constraint
^8.16.0 || ^9.0.0 - Data streams: Both logs and metrics properly configured with correct input types (
logfile,prometheus/metrics) - Ingest pipelines: Comprehensive handling of both JSON and logfmt log formats, proper Prometheus metric renaming
- Field definitions: Fields are properly typed with metric types (gauge/counter) and units where applicable
- Pipeline tests: Coverage across Grafana versions 10, 11, and 12, plus both log formats
- Time series support: Metrics manifest correctly enables
index_mode: time_series - Dashboards: Two dashboards covering overview metrics and log analysis
- Documentation: README with setup instructions and field reference tables
- Icons/Screenshots: Logo SVG and dashboard screenshots included
- Docker-compose: Proper test setup with healthcheck
🔬 Verification
elastic-package lint✅ passeselastic-package build✅ succeeds
💡 Optional Improvements
- The logs dashboard "eror" query filter (
log.level: eror) handles Grafana's unique spelling, but consider normalizingeror→errorin the ingest pipeline for consistency in searches - The
stat.data_keys.totalfield is defined in fields.yml but not covered in pipeline test fixtures
Overall this is a quality integration ready for merge after addressing the changelog PR link.
Why is Claude responding? | Type @claude to interact further
Give us feedback! React with 🚀 if perfect, 👍 if helpful, 👎 if not.
…link The logs dashboard now references the standard logs-* data view instead of the integration-specific logs-grafana.logs-* view. This resolves "field not found" warnings on Lens panels using ECS fields like user.name and http.request.method, since logs-* has full ECS field coverage. The event.dataset: grafana.logs filter already scopes results to Grafana logs. The overview dashboard's Datasources stat panel uses a formula to aggregate across labeled Prometheus metric documents. Also fixes the changelog PR link to reference elastic#17433.
…elog Major changes synced from local testing environment: Logs pipeline now extracts ECS fields from Grafana's structured log output — user.name, http.request.method, http.response.status_code, client.ip, url.path, event.duration, and log.logger are all mapped from their Grafana equivalents. Log level normalization handles Grafana's non-standard "eror" spelling. The logs dashboard references the standard logs-* data view instead of logs-grafana.logs-* to avoid "field not found" warnings on ECS fields. Metrics manifest enables TSDS (index_mode: time_series) with dimension fields for prometheus.labels.*, service.address, and service.name. The metrics pipeline includes the full _embedded_ecs dynamic template block. The overview dashboard's Datasources panel uses a formula to handle labeled Prometheus metric documents. Also adds correctly-named screenshot images referenced by the manifest, and fixes the changelog PR link to elastic#17433.
51f8e1f to
fbc55ad
Compare
The ecs@mappings component template is automatically composed into Fleet-managed index templates and already handles all ECS field type mappings. The inline _embedded_ecs block was duplicating ~460 lines of dynamic templates in each data stream manifest for no benefit.
|
Recent changes:
Ready for another test. |
Vale Linting ResultsSummary: 4 warnings, 20 suggestions found
|
| File | Line | Rule | Message |
|---|---|---|---|
| packages/grafana/docs/README.md | 380 | Elastic.Latinisms | Latin terms and abbreviations are a common source of confusion. Use 'for example' instead of 'e.g'. |
| packages/grafana/docs/README.md | 380 | Elastic.DontUse | Don't use 'please'. |
| packages/grafana/docs/README.md | 381 | Elastic.Latinisms | Latin terms and abbreviations are a common source of confusion. Use 'for example' instead of 'e.g'. |
| packages/grafana/docs/README.md | 394 | Elastic.Latinisms | Latin terms and abbreviations are a common source of confusion. Use 'for example' instead of 'e.g'. |
💡 Suggestions (20)
| File | Line | Rule | Message |
|---|---|---|---|
| packages/grafana/docs/README.md | 149 | Elastic.WordChoice | Consider using 'can, might' instead of 'may', unless the term is in the UI. |
| packages/grafana/docs/README.md | 152 | Elastic.WordChoice | Consider using 'can, might' instead of 'may', unless the term is in the UI. |
| packages/grafana/docs/README.md | 152 | Elastic.WordChoice | Consider using 'can, might' instead of 'may', unless the term is in the UI. |
| packages/grafana/docs/README.md | 210 | Elastic.WordChoice | Consider using 'start, run' instead of 'boot', unless the term is in the UI. |
| packages/grafana/docs/README.md | 211 | Elastic.WordChoice | Consider using 'start, run' instead of 'boot', unless the term is in the UI. |
| packages/grafana/docs/README.md | 212 | Elastic.WordChoice | Consider using 'start, run' instead of 'boot', unless the term is in the UI. |
| packages/grafana/docs/README.md | 212 | Elastic.WordChoice | Consider using 'start, run' instead of 'boot', unless the term is in the UI. |
| packages/grafana/docs/README.md | 213 | Elastic.WordChoice | Consider using 'start, run' instead of 'boot', unless the term is in the UI. |
| packages/grafana/docs/README.md | 213 | Elastic.WordChoice | Consider using 'start, run' instead of 'boot', unless the term is in the UI. |
| packages/grafana/docs/README.md | 228 | Elastic.Wordiness | Consider using 'because' instead of 'since'. |
| packages/grafana/docs/README.md | 259 | Elastic.Wordiness | Consider using 'because' instead of 'since'. |
| packages/grafana/docs/README.md | 356 | Elastic.Wordiness | Consider using 'tell' instead of 'inform'. |
| packages/grafana/docs/README.md | 374 | Elastic.WordChoice | Consider using 'can, might' instead of 'may', unless the term is in the UI. |
| packages/grafana/docs/README.md | 378 | Elastic.Wordiness | Consider using 'tell' instead of 'inform'. |
| packages/grafana/docs/README.md | 378 | Elastic.WordChoice | Consider using 'can, might' instead of 'may', unless the term is in the UI. |
| packages/grafana/docs/README.md | 378 | Elastic.WordChoice | Consider using 'can, might' instead of 'may', unless the term is in the UI. |
| packages/grafana/docs/README.md | 380 | Elastic.WordChoice | Consider using 'can, might' instead of 'may', unless the term is in the UI. |
| packages/grafana/docs/README.md | 380 | Elastic.WordChoice | Consider using 'deactivated, deselected, hidden, turned off, unavailable' instead of 'disabled', unless the term is in the UI. |
| packages/grafana/docs/README.md | 381 | Elastic.WordChoice | Consider using 'can, might' instead of 'may', unless the term is in the UI. |
| packages/grafana/docs/README.md | 394 | Elastic.WordChoice | Consider using 'can, might' instead of 'may', unless the term is in the UI. |
The Vale linter checks documentation changes against the Elastic Docs style guide.
To use Vale locally or report issues, refer to Elastic style guide for Vale.
| - src: /img/screenshot-overview.png | ||
| title: Grafana Overview Dashboard | ||
| size: 1024x1100 | ||
| type: image/png | ||
| - src: /img/screenshot-logs.png | ||
| title: Grafana Logs Dashboard | ||
| size: 1024x1100 | ||
| type: image/png |
There was a problem hiding this comment.
| - src: /img/screenshot-overview.png | |
| title: Grafana Overview Dashboard | |
| size: 1024x1100 | |
| type: image/png | |
| - src: /img/screenshot-logs.png | |
| title: Grafana Logs Dashboard | |
| size: 1024x1100 | |
| type: image/png | |
| - src: /img/grafana-overview.png | |
| title: Grafana Overview Dashboard | |
| size: 1024x1100 | |
| type: image/png | |
| - src: /img/grafana-logs.png | |
| title: Grafana Logs Dashboard | |
| size: 1024x1100 | |
| type: image/png |
Looks like these two were used accidentally, I see there are newer screenshots - better looking and with actual data.
Shall the screenshot-logs.png and screenshot-overview.png be deleted alltogether?
There was a problem hiding this comment.
Manually tested with Grafana 12.3.3 and Elastic Stack 9.2.4.
Overall good.
The following things need to be addressed:
- [Grafana] Overview: almost all Y-axis names simply copy the name of the first group/category. For example: panel "Memory" (see screenshot)
-
It would be helpful for the dashboards to contain a Link panel: https://www.elastic.co/docs/explore-analyze/visualize/link-panels
to be able to switch between dashboards quickly -
For new dashboards it is now desirable to have a markdown panel in the top - providing a high-level description of the dashboard
Minor stuff to address:
- Screenshots
- Vale linter warnings
| - set: | ||
| tag: set_event_kind | ||
| field: event.kind | ||
| value: event | ||
| - set: | ||
| tag: set_event_module | ||
| field: event.module | ||
| value: grafana | ||
| - set: | ||
| tag: set_event_dataset | ||
| field: event.dataset | ||
| value: grafana.logs |
There was a problem hiding this comment.
These fields are already defined in the base-fields.yml
- name: event.module
type: constant_keyword
value: grafana
description: Event module.
- name: event.dataset
type: constant_keyword
value: grafana.logs
description: Event dataset.
|
/test |
|
/test |
|
/test |
|
/test |
🚀 Benchmarks reportTo see the full report comment with |
2123a52 to
33bcae5
Compare
Dashboard UX fixes, pipeline cleanup, docs linting, screenshot renames. Also fix event.module rejection in TSDB by removing the value the promtheus input sets before indexing.
33bcae5 to
554eba2
Compare
|
Added changes as I think they were requested. |
|
/test |
Remove event.module and event.dataset from expected output since the pipeline no longer sets these constant_keyword fields.
5605043 to
9f12ed4
Compare
|
Sorry, will run the elstic-package tests myself before pushing in the future, I could've seen this myself. |
|
/test |
| | event.original | Raw text message of entire event. Used to demonstrate log integrity or where the full log message (before splitting it up in multiple parts) may be required, e.g. for reindex. This field is not indexed and doc_values are disabled. It cannot be searched, but it can be retrieved from `_source`. If users wish to override this and index this field, please see `Field data types` in the `Elasticsearch Reference`. | keyword | | ||
| | event.severity | The numeric severity of the event according to your event source. What the different severity values mean can be different between sources and use cases. It's up to the implementer to make sure severities are consistent across events from the same source. The Syslog severity belongs in `log.syslog.severity.code`. `event.severity` is meant to represent the severity according to the event source (e.g. firewall, IDS). If the event source does not publish its own severity, you may optionally copy the `log.syslog.severity.code` to `event.severity`. | long | | ||
| | grafana.log.handler | Request handler name. | keyword | | ||
| | grafana.log.orgId | Organization ID in context. | long | | ||
| | grafana.log.subUrl | Grafana sub-URL prefix. | keyword | | ||
| | host.containerized | Whether the host is a container. | boolean | | ||
| | host.name | Host name. | keyword | | ||
| | host.os.build | OS build information. | keyword | | ||
| | host.os.codename | OS codename, if any. | keyword | | ||
| | http.request.method | HTTP request method. The value should retain its casing from the original event. For example, `GET`, `get`, and `GeT` are all considered valid values for this field. | keyword | | ||
| | http.request.referrer | Referrer for this HTTP request. | keyword | | ||
| | http.response.body.bytes | Size in bytes of the response body. | long | | ||
| | http.response.status_code | HTTP response status code. | long | | ||
| | input.type | Type of Filebeat input. | keyword | | ||
| | log.level | Original log level of the log event. If the source of the event provides a log level or textual severity, this is the one that goes in `log.level`. If your source doesn't specify one, you may put your event transport's severity here (e.g. Syslog severity). Some examples are `warn`, `err`, `i`, `informational`. | keyword | |
There was a problem hiding this comment.
Please address these Vale lint warning
mykola-elastic
left a comment
There was a problem hiding this comment.
LGTM
Please, address Vale lint warnings and please try removing user: "0" from docker-compose
|
In the latest commit, I also removed the |
|
/test |
💚 Build Succeeded
History
|
lalit-satapathy
left a comment
There was a problem hiding this comment.
Code owner approval for merge
|
Package grafana - 0.1.0 containing this change is available at https://epr.elastic.co/package/grafana/0.1.0/ |
* Add Grafana integration New integration for collecting logs and Prometheus metrics from Grafana instances (10.x, 11.x, 12.x). Includes ingest pipelines for both JSON and logfmt log formats, field mappings that rename prometheus.metrics.* to a structured grafana.* namespace, two Kibana dashboards (overview and logs), and pipeline tests with data captured from a real Grafana 12.3.2 instance. * Switch logs dashboard to logs-*, add ECS field mapping, and fix changelog Major changes synced from local testing environment: Logs pipeline now extracts ECS fields from Grafana's structured log output — user.name, http.request.method, http.response.status_code, client.ip, url.path, event.duration, and log.logger are all mapped from their Grafana equivalents. Log level normalization handles Grafana's non-standard "eror" spelling. The logs dashboard references the standard logs-* data view instead of logs-grafana.logs-* to avoid "field not found" warnings on ECS fields. Metrics manifest enables TSDS (index_mode: time_series) with dimension fields for prometheus.labels.*, service.address, and service.name. The metrics pipeline includes the full _embedded_ecs dynamic template block. The overview dashboard's Datasources panel uses a formula to handle labeled Prometheus metric documents. Also adds correctly-named screenshot images referenced by the manifest, and fixes the changelog PR link to #17433. * Remove redundant _embedded_ecs dynamic templates from manifests The ecs@mappings component template is automatically composed into Fleet-managed index templates and already handles all ECS field type mappings. The inline _embedded_ecs block was duplicating ~460 lines of dynamic templates in each data stream manifest for no benefit. * Update CODEOWNERS * Update docs/README.md: run `elastic-package build` to regenerate it * Update test-json-log.json-expected.json * Update test-logfmt-log.log-expected.json * Fix PR review feedback and TSDB event.module issue Dashboard UX fixes, pipeline cleanup, docs linting, screenshot renames. Also fix event.module rejection in TSDB by removing the value the promtheus input sets before indexing. * Update pipeline test expected output Remove event.module and event.dataset from expected output since the pipeline no longer sets these constant_keyword fields. * Add requested changes * Fix CODEOWNERS: remove unrelated trailing newline diff * Remove unnecessary import_mappings from build.yml --------- Co-authored-by: Mykola Kmet <mykola.kmet@elastic.co>
* Add Grafana integration New integration for collecting logs and Prometheus metrics from Grafana instances (10.x, 11.x, 12.x). Includes ingest pipelines for both JSON and logfmt log formats, field mappings that rename prometheus.metrics.* to a structured grafana.* namespace, two Kibana dashboards (overview and logs), and pipeline tests with data captured from a real Grafana 12.3.2 instance. * Switch logs dashboard to logs-*, add ECS field mapping, and fix changelog Major changes synced from local testing environment: Logs pipeline now extracts ECS fields from Grafana's structured log output — user.name, http.request.method, http.response.status_code, client.ip, url.path, event.duration, and log.logger are all mapped from their Grafana equivalents. Log level normalization handles Grafana's non-standard "eror" spelling. The logs dashboard references the standard logs-* data view instead of logs-grafana.logs-* to avoid "field not found" warnings on ECS fields. Metrics manifest enables TSDS (index_mode: time_series) with dimension fields for prometheus.labels.*, service.address, and service.name. The metrics pipeline includes the full _embedded_ecs dynamic template block. The overview dashboard's Datasources panel uses a formula to handle labeled Prometheus metric documents. Also adds correctly-named screenshot images referenced by the manifest, and fixes the changelog PR link to elastic#17433. * Remove redundant _embedded_ecs dynamic templates from manifests The ecs@mappings component template is automatically composed into Fleet-managed index templates and already handles all ECS field type mappings. The inline _embedded_ecs block was duplicating ~460 lines of dynamic templates in each data stream manifest for no benefit. * Update CODEOWNERS * Update docs/README.md: run `elastic-package build` to regenerate it * Update test-json-log.json-expected.json * Update test-logfmt-log.log-expected.json * Fix PR review feedback and TSDB event.module issue Dashboard UX fixes, pipeline cleanup, docs linting, screenshot renames. Also fix event.module rejection in TSDB by removing the value the promtheus input sets before indexing. * Update pipeline test expected output Remove event.module and event.dataset from expected output since the pipeline no longer sets these constant_keyword fields. * Add requested changes * Fix CODEOWNERS: remove unrelated trailing newline diff * Remove unnecessary import_mappings from build.yml --------- Co-authored-by: Mykola Kmet <mykola.kmet@elastic.co>
New integration for collecting logs and Prometheus metrics from Grafana instances (10.x, 11.x, 12.x). Includes ingest pipelines for both JSON and logfmt log formats, field mappings that rename prometheus.metrics.* to a structured grafana.* namespace, two Kibana dashboards (overview and logs), and pipeline tests with data captured from a real Grafana 12.3.2 instance.
Proposed commit message
This adds a new
grafanaintegration package with two data streams: logs and metrics.The logs data stream reads Grafana server log files. The ingest pipeline handles both JSON and logfmt formats — it tries JSON parsing first and falls back to logfmt via grok. Both formats are supported across Grafana 10, 11, and 12 since the field names are identical across those versions. The pipeline extracts HTTP request fields (method, path, status, duration, remote_addr), maps log levels to ECS severity values, and populates the
grafana.log.*namespace.The metrics data stream scrapes Grafana's Prometheus
/metricsendpoint using theprometheus/metricsinput type. The ingest pipeline renamesprometheus.metrics.*fields into a structuredgrafana.*namespace — instance stats go tografana.stat.*, database connection pool metrics tografana.database.connections.*, Go runtime tografana.go.*, and so on. There are explicit renames for about 100 known metrics, plus a catch-all painless script that handles any remaining metrics by convention.Two Kibana dashboards are included: an overview dashboard showing instance stats, resource usage, database connections, and alerting status, and a logs dashboard with log volume over time, level breakdown, top error messages, HTTP status codes, and a log stream panel.
Pipeline test fixtures use real data captured from a Grafana 12.3.2 instance. The metrics tests cover three versions (v10, v11, v12) with progressively more metrics present in each, reflecting the additive nature of Grafana's metric surface across major versions.
Checklist
changelog.ymlfile.Author's Checklist
elastic-package lintpasseselastic-package buildsucceedselastic-package test pipeline)How to test this PR locally
For a full end-to-end test with a running Grafana instance:
elastic-package stack up -v -d elastic-package test systemOr point the agent at any Grafana instance with metrics enabled (
GF_METRICS_ENABLED=true) and a log file path configured. Both logfmt (default) and JSON (GF_LOG_FILE_FORMAT=json) log formats are handled automatically.Related issues
Screenshots
[Grafana] Overview Dashboard

[Grafana] Logs Dashboard
