Skip to content

Adding Pantherflow Baseline Builders - #2013

Merged
zaynahsmith-dasilva merged 10 commits into
developfrom
pantherflow_okta
Apr 9, 2026
Merged

zaynahsmith-dasilva merged 10 commits into
developfrom
pantherflow_okta

Conversation

@zaynahsmith-dasilva

Copy link
Copy Markdown
Contributor

Background

Changes

  • The Pantherflow equivalents of the Okta AD Baseline Builders were added to the Lookup Tables.

Testing

@zaynahsmith-dasilva
zaynahsmith-dasilva requested a review from a team as a code owner April 6, 2026 17:50
@cursor

cursor Bot commented Apr 6, 2026 •

Copy link
Copy Markdown

PR Summary

Medium Risk
Renames and replaces Okta AD Agent baseline lookup tables and updates the z-score anomaly query/rule to point at the new table name, which could break detections if deployments still reference the old lookups or schedules.

Overview
Migrates the Okta AD Agent baseline builders (30/90/180 day) from the legacy okta_ad_agent_baseline_* lookup tables to new PantherFlow-backed lookup tables okta_ad_pantherflow_baseline_*, including adding LogTypeMap/refresh configuration and removing the old SQL-based lookup table definitions.

Updates the Okta pack and the AD-agent z-score anomaly query/rule docs to reference the new 90-day lookup table, and adjusts generated indexes/coverage metadata and deprecated.txt to reflect the deprecation of the prior baseline builders and lookup names.

Reviewed by Cursor Bugbot for commit 09953fc. Bugbot is set up for automated code reviews on this repo. Configure here.

@zaynahsmith-dasilva zaynahsmith-dasilva added the rules Real-time log data detections label Apr 6, 2026
@github-actions

github-actions Bot commented Apr 6, 2026

Copy link
Copy Markdown

📐 Style Guide Review

Analyzing commit: 105c759aa40b513a3c2c449db710f024bbb4e2c4

Style Compliance: EXCELLENT

📋 Style Findings

✅ Compliant Areas

  • YAML Structure: All files follow proper 2-space indentation and YAML formatting conventions
  • Required Metadata: All scheduled queries include required fields (AnalysisType, QueryName, Enabled, Description, Query, Schedule, Tags)
  • Query Formatting: SQL/KQL queries are well-formatted with consistent indentation, clear variable names, and proper use of PantherFlow pragma

⚠️ Style Issues

  • Pack Reference Consistency: Pack file (packs/okta.yml:46-51) references both descriptive QueryNames and separate lookup table IDs ("okta_baseline_30d", etc.) that may not correspond to actual query definitions
  • Comment Formatting: SQL queries could benefit from more inline comments explaining complex aggregation logic, especially around z-score baseline calculations

📝 Recommendations

  • Query Documentation: Consider adding inline comments within the complex JOIN and aggregation sections to improve maintainability
  • Naming Alignment: Verify that pack references align with intended lookup table population strategy

The changes demonstrate excellent adherence to YAML formatting standards and Panther detection conventions. The scheduled queries are well-structured with appropriate metadata and consistent implementation patterns across all three time windows (30d, 90d, 180d).

@github-actions

github-actions Bot commented Apr 6, 2026

Copy link
Copy Markdown

🐍 Python Logic Review

Analyzing commit: 105c759aa40b513a3c2c449db710f024bbb4e2c4

Logic Quality: N/A - No Python Files

🔍 Logic Analysis

📋 Review Summary

This PR contains no Python (.py) files to review. All changed files are YAML configuration files:

  • packs/okta.yml - Pack configuration listing detection IDs
  • lookup_tables/okta/okta_ad_pantherflow_baseline_*.yml - Scheduled queries using PantherFlow (not Python)

✅ Files Reviewed

  • Pack Configuration: Well-structured YAML pack with proper detection ID references
  • Scheduled Queries: PantherFlow queries for behavioral baseline generation (30/90/180 day windows)

📝 Note

No Python logic review needed as this PR only contains YAML configuration and query definitions.

@github-actions

github-actions Bot commented Apr 6, 2026

Copy link
Copy Markdown

📄 YAML Metadata Review - Analyzing commit: 105c759 - Metadata Quality: GOOD - Well-Documented: Comprehensive descriptions in scheduled queries explaining behavioral baseline use case, appropriate tags, correct AnalysisType - Issues: Pack description too brief, missing References fields for PantherFlow docs, QueryNames could be shorter for UI - Suggestions: Add References to PantherFlow documentation, expand pack description to mention key security use cases covered

@github-actions

github-actions Bot commented Apr 6, 2026

Copy link
Copy Markdown

🧪 Testing Review

Analyzing commit: 105c759aa40b513a3c2c449db710f024bbb4e2c4

Test Quality: NOT APPLICABLE

🔬 Test Analysis

ℹ️ No Test Cases Required

The changed files in this PR do not require traditional unit test cases:

  • packs/okta.yml - Pack definition file (AnalysisType: pack) that groups detection IDs. Packs don't contain detection logic or test cases.
  • lookup_tables/okta/okta_ad_pantherflow_baseline_*.yml - Scheduled queries (AnalysisType: scheduled_query) that generate behavioral baseline data. These don't follow the rule/policy testing pattern.

✅ Appropriate File Types

  • All files are correctly structured for their respective types
  • Scheduled queries include proper metadata (QueryName, Description, Schedule, Tags)
  • Pack file correctly lists detection IDs and dependencies
  • No traditional Tests: sections needed for these analysis types

📝 Documentation Quality

  • 180-day baseline: Clear description of 180-day behavioral baseline computation
  • 90-day baseline: Well-documented 90-day window analysis
  • 30-day baseline: Consistent documentation pattern across all baseline queries
  • All queries properly tagged with relevant categories (Okta, Active Directory, Baseline, etc.)

Note: These files generate baseline data for behavioral analytics rather than performing detection logic, so they don't require the standard unit test format used for rules and policies.

@github-actions

github-actions Bot commented Apr 6, 2026

Copy link
Copy Markdown

🔍 Query Review

Analyzing commit: 105c759aa40b513a3c2c449db710f024bbb4e2c4

Query Quality: GOOD

📊 Query Analysis

✅ Good Practices

  • Excellent time filtering: All queries use proper time ranges with p_event_time and smart exclusion of last 7 days to avoid incomplete data
  • Well-structured KQL: Clean use of PantherFlow syntax with logical let statements and clear data flow
  • Appropriate aggregations: Statistical calculations (mean, stddev) are properly rounded and meaningful for behavioral baselines

⚠️ Query Issues

  • Filename format: Files don't follow standard <LogProvider>.<QueryTitle>.Query.yml convention (though they're in lookup_tables, not queries/)
  • Display name inconsistency: Uses QueryName instead of DisplayName field
  • Schedule consistency: All three queries have identical 30-day schedules despite different baseline windows

🚀 Performance Suggestions

  • Large time windows: 180-day lookbacks with complex aggregations could be expensive - monitor execution times
  • Filtering efficiency: The total_events >= 10 threshold is good for reducing noise and improving performance
Comment thread packs/okta.yml Outdated
@github-actions

github-actions Bot commented Apr 6, 2026

Copy link
Copy Markdown

🏗️ Architecture Compliance Review

Analyzing commit: 105c759aa40b513a3c2c449db710f024bbb4e2c4

Architecture: GOOD

🔧 Architecture Analysis

✅ Architectural Strengths

  • Proper File Organization: Scheduled queries correctly placed in /lookup_tables/okta/ directory and pack appropriately updated in /packs/
  • Consistent Query Structure: All three baseline queries (30d, 90d, 180d) follow identical patterns with only time window variations, promoting maintainability
  • Appropriate Detection Type: Uses AnalysisType: scheduled_query correctly for behavioral baseline generation rather than dual-file rule/policy architecture

⚠️ Architecture Issues

  • Naming Inconsistency: Pack references both descriptive names ("Okta 30 Day AD Pantherflow in Data Explorer") and short names ("okta_baseline_30d") - unclear if these map to the same resources or if lookup table definitions are missing
  • Schedule Uniformity: All queries use identical 30-day schedules (43200 minutes) regardless of their baseline window (30d/90d/180d) - may need clarification on intended refresh rates

🎯 Architecture Recommendations

  • Clarify Lookup Table Mapping: Ensure pack references align with actual scheduled query names or add missing lookup table definitions for the short-name references
  • Consider Schedule Optimization: Evaluate if different baseline windows should have different refresh frequencies based on their analytical purpose

Review Complete ✅
All review sections have been posted. Check the comments above for detailed findings in each area.

Comment thread lookup_tables/okta/okta_ad_pantherflow_baseline_90day.yml
Comment thread deprecated.txt Outdated

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit c7bcda7. Configure here.

Comment thread lookup_tables/okta/okta_ad_pantherflow_baseline_90day.yml Outdated
@zaynahsmith-dasilva
zaynahsmith-dasilva added this pull request to the merge queue Apr 9, 2026
Merged via the queue into develop with commit 3e7da03 Apr 9, 2026
18 checks passed
@zaynahsmith-dasilva
zaynahsmith-dasilva deleted the pantherflow_okta branch April 9, 2026 18:28
zaynahsmith-dasilva added a commit that referenced this pull request Apr 10, 2026
Keep AWS.S3.Disable.Security.Controls from branch and add Okta AD Pantherflow
baseline entries added in develop (#2013).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

rules Real-time log data detections

3 participants