[Elasticsearch][Ingest pipeline] Stop truncating the elasticsearch server log messages - #12813
Merged
Conversation
Contributor
🚀 Benchmarks reportTo see the full report comment with |
|
💚 Build Succeeded
History
cc @consulthys |
Contributor
|
Package elasticsearch - 1.17.2 containing this change is available at https://epr.elastic.co/package/elasticsearch/1.17.2/ |
flexitrev
pushed a commit
that referenced
this pull request
Mar 20, 2025
…rver log messages (#12813) * Stop truncating the elasticsearch server log messages * Add more test cases
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.




Proposed commit message
This PR fixes the ingest pipeline called
logs-elasticsearch.server-<version>-pipeline-jsonthat parses elasticsearch server logs.The pipeline was inherited from the Filebeat
elasticsearchmodule and hasn't changed in several years. The main issue is that the pipeline makes the assumption that if themessagefield value starts with square brackets (e.g.[xyz] some log message), thenxyzis considered to be an index name (indexed in theelasticsearch.index.namefield) and the message is truncated to only what comes after the square brackets (i.e.some log message). This assumption might have been true at some point in the past, but isn't the case anymore, i.e. the square brackets can contain literally anything, such as component names, class names, etc. Truncating themessagefield breaks downstream processes that expect to find the full log message in that field.For instance, when applied on the following log message
the truncated
messagefield will only containis not ready for collect yet, which now lacks context and is unsuable.As it is not easy to find out all (internal and external) downstream processes that rely on the index name to be extracted from the log message, we need to keep extracting whatever is in the square brackets, but without truncating the
messagefield. This PR suggests a non-breaking change that will keep extracting the index name (if it doesn't already exist in the document), while leaving themessagefield alone.Checklist
changelog.ymlfile.How to test this PR locally
messagefield is unalteredResponse:
If the
elasticsearch.index.namefield is already present in the log document, then it will not be overriddenResponse:
Related issues
Closes #12501