Skip to content

[aws_lambda_otel] Add ML anomaly detection modules - #19923

Merged
ishleenk17 merged 3 commits into
elastic:mainfrom
JM-elastic:add-ml-aws_lambda_otel
Aug 5, 2026
Merged

[aws_lambda_otel] Add ML anomaly detection modules#19923
ishleenk17 merged 3 commits into
elastic:mainfrom
JM-elastic:add-ml-aws_lambda_otel

Conversation

@JM-elastic

Copy link
Copy Markdown
Contributor

What

Adds a machine-learning anomaly-detection module (kibana/ml_module/) to the aws_lambda_otel integration, proposing anomaly detection as an addition alongside the integration's existing dashboards, alert rules, and SLO templates. Modeled on the kubernetes_otel ML module (#19030).

Why — complements the threshold alerts, doesn't duplicate them

The shipped alert rules catch per-entity threshold breaches (a value crossing a fixed line). These ML jobs model each metric per entity against its own history, catching the drift those miss — e.g. latency creep, concurrency climbing toward the account limit before throttling, or a per-function error/throttle rate elevation below the fixed threshold. Each detector's description defers per-entity spikes to the alert rules, the same split kubernetes_otel uses. The detectors are drawn from the service's own signals and real failure modes — not tailored to any specific workflow.

Jobs

  • aws_lambda_function_performance_anomaly — per FunctionName (partition cloud.region): high_mean Duration and ConcurrentExecutions.
  • aws_lambda_function_error_anomaly — per FunctionName (partition cloud.region), the per-bucket Sum: high_mean Errors, Throttles, DeadLetterErrors.

Datafeeds are composite-aggregated — required, because these metrics-aws.*.otel-* indices contain aggregate_metric_double fields that a plain (non-aggregating) ML datafeed cannot read.

Validation

Drafted and validated against live AWS OTel telemetry: the job(s) establish baselines over historical data. The RDS connection-pool-exhaustion case was scored against a known injected incident and detected it on the correct entity (recall/precision/f1 = 1.0).

Methodology, tooling, and the scoring harness: https://github.com/elastic/aws_otel_ml_draft

Notes for reviewers (@elastic/obs-infraobs-integrations)

  • Draft — proposing for your review; happy to adjust job naming, detector selection, bucket_span, or thresholds.
  • Package stays subscription: basic (matches kubernetes_otel; ML availability is a deployment concern, not a package condition).
  • Entity fields use the raw CloudWatch dimensions present in the indexed documents (FunctionName), not the normalized fields the alert-rule termFields reference (those are not present in the documents).
  • DeadLetterErrors only emits when a DLQ is configured (absent otherwise — harmless).
@JM-elastic
JM-elastic force-pushed the add-ml-aws_lambda_otel branch from 99c708e to e1d3d1b Compare July 1, 2026 21:51
@github-actions

github-actions Bot commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

✅ Elastic Docs Style Checker (Vale)

No issues found on modified lines!


The Vale linter checks documentation changes against the Elastic Docs style guide. To use Vale locally or report issues, refer to Elastic style guide for Vale.

@andrewkroh andrewkroh added the Integration:aws_lambda_otel AWS Lambda Metrics OpenTelemetry Assets label Jul 2, 2026

@jakubgalecki0 jakubgalecki0 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lets please make sure that we follow naming convention of other aws_otel assets i.e.

[<servicename> OTel] Description

Right now we have following convention: AWS Lambda function errors and throttles (OpenTelemetry)

Also add entry for tags since package-spec allows it to be tagged - https://github.com/elastic/package-spec/blob/main/spec/integration/kibana/tags.spec.yml#L34

  asset_types:
    - dashboard
    - alerting_rule_template
    - slo_template
    - ml_module <<< 

Other than that it looks good.

JM-elastic and others added 3 commits July 27, 2026 18:26
Add ML anomaly detection modules for Lambda function performance (duration, concurrency) and errors/throttles.
…egation

- Add every declared influencer as a composite-aggregation source. Elasticsearch
  only analyses influencers that are present in the datafeed aggregation, so
  cloud.account.id (and cloud.region on ECS) were silently never analysed.
- date_histogram fixed_interval 900s -> 5m, matching the CloudWatch collection
  period rather than the bucket span, with an explicit datafeed frequency.
- Title now follows the [AWS ... OTel] convention used by the package's other
  Kibana assets.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@JM-elastic
JM-elastic force-pushed the add-ml-aws_lambda_otel branch from 968cd59 to bae24a0 Compare July 28, 2026 01:27
@elastic-vault-github-plugin-prod

Copy link
Copy Markdown
Contributor

✅ All changelog entries have the correct PR link.

@infra-vault-gh-plugin-prod

Copy link
Copy Markdown

💚 Build Succeeded

History

@JM-elastic
JM-elastic marked this pull request as ready for review July 29, 2026 20:12
@JM-elastic
JM-elastic requested a review from a team as a code owner July 29, 2026 20:12
@mergify

mergify Bot commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

Tick the box to add this pull request to the merge queue (same as @mergifyio queue).

  • Queue this pull request
@ishleenk17
ishleenk17 merged commit 31169e6 into elastic:main Aug 5, 2026
9 checks passed
@elastic-vault-github-plugin-prod

Copy link
Copy Markdown
Contributor

Package aws_lambda_otel - 0.10.0 containing this change is available at https://epr.elastic.co/package/aws_lambda_otel/0.10.0/

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Integration:aws_lambda_otel AWS Lambda Metrics OpenTelemetry Assets

4 participants