[Check Point] Process the packets field in SecureXL format - #16235
Conversation
🚀 Benchmarks reportTo see the full report comment with |
P1llus
left a comment
There was a problem hiding this comment.
Do we have any sample data (anonymized) we can add to the pipeline test?
| field: checkpoint.subs_exp | ||
| ignore_missing: true | ||
| - script: | ||
| tag: script_parse_checkpoint_packets_dropped |
There was a problem hiding this comment.
It would be preferred if we could use a common processor to parse this out, as script regex is quite heavy, was this the only way? looking at the data I guess it does seem like that, but its quite troublesome for performance reasons
There was a problem hiding this comment.
I welcome any ideas 😄. I tried split + grok but couldn't make it work.
There was a problem hiding this comment.
Yeah once I noticed it was meant to produce an array then any other processor mostly goes out the window with the exception of maybe foreach, but that just introduces many other issues, so script was the way to go, I was just hoping we could have managed to skip the regex part, as in theory the data is structured in some ways.
There was a problem hiding this comment.
I see. Yeah a review of https://www.elastic.co/docs/reference/scripting-languages/painless/painless-ingest-processor-context has shown to me how this is possible, so I've implemented a non-regexp version.
| description: | | ||
| Amount of packets dropped. | ||
| - name: packets_dropped | ||
| type: nested |
There was a problem hiding this comment.
Just leaving it as a note rather than something that needs to be resolved, nested types are something we do try to avoid as much as possible, but at this point I am unsure about any other better alternative
There was a problem hiding this comment.
Yeah it's a tradeoff. With nested types we accurately show the underlying structure, so that they can do "find packets where source is X and destination is Y", but it does come at a cost. The alternative could be to use a group but then we'll only have a list of all sources and a list of all destinations in the event and the information about which source goes to which destination would be lost.
There was a problem hiding this comment.
I don't know the details, but my assumption was that there is no performance penalty as long as nobody is searching on that field.
There was a problem hiding this comment.
Yeah nested fields don't induce a performance penalty but you cannot use aggregations on them, so any dashboard visualization etc wont work, the only things you can do is querying for example on the interface name of any of the objects
| - script: | ||
| tag: script_parse_checkpoint_packets_dropped | ||
| description: Parse packets field containing connection tuples into structured packets_dropped array. | ||
| if: ctx.checkpoint?.packets instanceof String && ctx.checkpoint.packets.contains('<') |
There was a problem hiding this comment.
I am a bit unsure but it could be startsWith is better performance wise, though if the only times this is a string is when it is the SecureXL format it might be okay.
If we had sample data we could see if there was an already parsed field that showcases that it's SecureXL type data so we wouldn't need to use either
There was a problem hiding this comment.
Good point. I don't know if their tuple format is well-documented, for example if it's guaranteed that it does not start with a whitespace. The cleanest way is probably to try a conversion to number first and only run the script if it fails – I'll do that.
There was a problem hiding this comment.
It actually starts with a whitespace – < – so I went with trim().startsWith('<')
|
Hi! We just realized that we haven't looked into this PR in a while. We're sorry! We're labeling this issue as |
|
Pinging @elastic/integration-experience (Team:Integration-Experience) |
💚 Build Succeeded
History
cc @ilyannn |
| - script: | ||
| tag: script_parse_checkpoint_packets_dropped | ||
| description: Parse packets field containing connection tuples into structured packets_dropped array. | ||
| if: ctx.checkpoint?.packets instanceof String && ctx.checkpoint.packets.trim().startsWith('<') |
There was a problem hiding this comment.
As a note, if statements are also painless scripts, so could move this into the source of this processor and it should work just the same, basically one script invocation instead of two. I've never benchmarked it (and not sure if it matters, I don't think conditionals are included in benchmarks), but it's probably very minimal overhead, especially once the JIT has warmed up.
There was a problem hiding this comment.
Sure but this way we have a clear visual on when the script is run, I think it's moire readable this way.
|
Package checkpoint - 1.45.0 containing this change is available at https://epr.elastic.co/package/checkpoint/1.45.0/ |
…6235) The specification at https://support.checkpoint.com/results/sk/sk144192 describes the `packets` field in the _Security Gateway - SecureXL Fields_ section as the tuple of - Source IP address - Source Port - Destination IP address - Destination Port - Protocol Number The current pipeline, however, parses the `packets` field as an integer, handling only the format described in the _Security Gateway - Firewall Fields_ section. When this field comes as defined in the SecureXL section – a string describing the dropped packets – the current pipeline fails. In practice the tuple can also contain an additional member denoting the network interface. This PR adds the parsing to the pipeline. The result is an array of json objects, each structured according to the ECS/ We handle both the case when the interface name is present (as in practice) and when it's not (as per specification). The added group is defined as `nested` to preserve the relationships within each tuple. The parsing is implemented as a Painless script; alternative approaches were tried, but were not ultimately successful. Cursor with the Claude Opus 4.5 model was used in creating this PR's contents, especially the Painless script.
Proposed commit message
The specification at https://support.checkpoint.com/results/sk/sk144192 describes the
packetsfield in the Security Gateway - SecureXL Fields section as the tuple ofThe current pipeline, however, parses the
packetsfield as an integer, handling only the format described in the Security Gateway - Firewall Fields section. When this field comes as defined in the SecureXL section – a string describing the dropped packets – the current pipeline fails.In practice the tuple can also contain an additional member denoting the network interface.
This PR adds the parsing to the pipeline. The result is an array of json objects, each structured according to the ECS, that is:
We handle both the case when the interface name is present (as in practice) and when it's not (as per specification). The added group is defined as
nestedto preserve the relationships within each tuple.The parsing is implemented as a Painless script; alternative approaches were tried, but were not ultimately successful.
Screenshot
From the specification:
Checklist
changelog.ymlfile.[ ] I have verified that any added dashboard complies with Kibana's Dashboard good practicesAuthor's Checklist
elastic-package checkHow to test this PR locally
GenAI Note
Cursor with the Claude Opus 4.5 model was used in creating this PR's contents, especially the Painless script.