Skip to content
Please update to the latest release 0.77.3 to address Multiple CVEs.
Server.Monitor.Client.Errors.Alert

Server.Monitor.Client.Errors.Alert

Create alerts from errors logged by client event queries.

Client event queries are often used for important monitoring and automation, but failures in these queries may otherwise go unnoticed in the monitoring logs. This artifact periodically inspects the monitoring logs for active client event artifacts and creates alerts for log entries that match IncludeFilter and are not rejected by ExcludeFilter.

Matching can be controlled by artifact name, log level, and message content. This allows you to alert on all client event query errors, or to focus on a smaller set of important artifacts and known failure patterns.

Each alert includes useful context such as the artifact name, log level, error message, and the original log timestamp. Duplicate alerts may be suppressed for a configurable interval to avoid repeated notifications for the same problem. Deduplication is performed on the alert name, which includes the event artifact name. This means that at most one alert is produced per artifact within the deduplication interval. You should check the event query log for errors; there may be more than one. Several clients may also be affected.

If SeverityField is set, the alert will also include a severity value taken from the matching IncludeFilter row’s Severity column. This can be used by Server.Monitor.Alerts to format, filter, and forward the alert as an e-mail notification, or by Server.Monitor.Alerts.UserMessage to post it as an in-app notification.

Any extra columns added to IncludeFilter are passed through to the alert under the same name as extra context. For instance, if you want to add a helpful description to a particular error, you can use IncludeFilter as follows:

Artifact Level Message Severity Explanation
.+ DEFAULT watch_ebpf: flags.PrepareFilterMapsFromPolicies: policy \S+ already exists? medium eBPF policy conflict (fixed in 0.76.1)
.+ DEFAULT fork/exec .+ no such file or directory high Executable not found on system
.+ ERROR .+ medium

With this filter we pick up any attempt to use execve with a non-existent binary. This would normally not produce an ERROR-level log entry (only a DEFAULT-level one), but it would definitely break important functionality in most artifacts. We also pick up eBPF policy name conflicts — a bug fixed in 0.76.1, which prevented eBPF monitoring from working in older versions when more than one eBPF-using artifact was active.

Remember that the filters are matched from top to bottom, so put your most specific filters at the top.

Note that many errors produced by native VQL functions and plugins are logged with the DEFAULT level rather than ERROR. To alert on these, either add your own checks that log ERROR to artifacts, or match the native log messages with IncludeFilter. See this reference for a list of regexes matching DEFAULT-level errors worth monitoring.

Note that this artifact periodically iterates over all clients in order to query their monitoring logs. The work is parallelized across clients and artifacts, but can still be expensive on large deployments.

#monitoring #alerts #notifications #errors


name: Server.Monitor.Client.Errors.Alert
author: Andreas Misje – @misje
description: |
  Create alerts from errors logged by client event queries.

  Client event queries are often used for important monitoring and automation,
  but failures in these queries may otherwise go unnoticed in the monitoring logs.
  This artifact periodically inspects the monitoring logs for active client event
  artifacts and creates alerts for log entries that match `IncludeFilter` and
  are not rejected by `ExcludeFilter`.

  Matching can be controlled by artifact name, log level, and message content.
  This allows you to alert on all client event query errors, or to focus on a
  smaller set of important artifacts and known failure patterns.

  Each alert includes useful context such as the artifact name, log level, error
  message, and the original log timestamp. Duplicate alerts may be suppressed for
  a configurable interval to avoid repeated notifications for the same problem.
  Deduplication is performed on the alert name, which includes the event artifact
  name. This means that at most one alert is produced per artifact within the
  deduplication interval. You should check the event query log for errors; there
  may be more than one. Several clients may also be affected.

  If `SeverityField` is set, the alert will also include a severity value taken
  from the matching `IncludeFilter` row's `Severity` column. This can be used
  by [`Server.Monitor.Alerts`](/exchange/artifacts/pages/server.monitor.alerts/)
  to format, filter, and forward the alert as an e-mail notification, or by
  [`Server.Monitor.Alerts.UserMessage`](/exchange/artifacts/pages/server.monitor.alerts.usermessage/)
  to post it as an in-app notification.

  Any extra columns added to `IncludeFilter` are passed through to the alert
  under the same name as extra context. For instance, if you want to add a
  helpful description to a particular error, you can use `IncludeFilter` as
  follows:

  | Artifact | Level | Message | Severity | Explanation |
  | -------- | ----- | ------- | -------- | ----------- |
  | .+ | DEFAULT | watch_ebpf: flags.PrepareFilterMapsFromPolicies: policy \S+ already exists? | medium | eBPF policy conflict (fixed in 0.76.1) |
  | .+ | DEFAULT | fork/exec .+ no such file or directory | high | Executable not found on system |
  | .+ | ERROR | .+ | medium | |

  With this filter we pick up any attempt to use [`execve`](/vql_reference/popular/execve/) with a non-existent
  binary. This would normally not produce an `ERROR`-level log entry (only a
  `DEFAULT`-level one), but it would definitely break important functionality
  in most artifacts. We also pick up eBPF policy name conflicts — a bug fixed
  in 0.76.1, which prevented eBPF monitoring from working in older versions
  when more than one eBPF-using artifact was active.

  Remember that the filters are matched from top to bottom, so put your most
  specific filters at the top.

  Note that many errors produced by native VQL functions and plugins
  are logged with the `DEFAULT` level rather than `ERROR`. To alert on
  these, either add your own checks that log `ERROR` to artifacts, or
  match the native log messages with `IncludeFilter`. See [this
  reference](/knowledge_base/tips/vql_error_catalogue/) for a list of
  regexes matching `DEFAULT`-level errors worth monitoring.

  Note that this artifact periodically iterates over **all** clients in
  order to query their monitoring logs. The work is parallelized across
  clients and artifacts, but can still be expensive on large deployments.

  #monitoring #alerts #notifications #errors

type: SERVER_EVENT

parameters:
  - name: Period
    type: int
    description: |
      Seconds between each monitoring log inspection
    default: 60

  - name: DedupInterval
    type: int
    description: |
      Suppress duplicate alerts from the same artifact within this many
      seconds. Inspect the event query log directly for the full set of
      errors, and note that several clients may be affected by the same issue.
    default: 3600

  - name: IncludeFilter
    type: csv
    description: |
      Include only log entries matching one of these rows. Each column is a
      regex; empty means "any". The optional `Severity` column is copied into
      the alert (see `SeverityField`). Any extra columns are passed through
      to the alert under the same name. Rows are evaluated top-to-bottom and
      the first match wins, so place specific rules before broad catch-alls.
    default: |
      Artifact,Level,Message,Severity
      .+,ERROR,.+,medium

  - name: ExcludeFilter
    type: csv
    description: |
      Reject log entries matching one of these rows, applied after
      `IncludeFilter`
    default: |
      Artifact,Level,Message
      Generic\.Client\.Stats$,.+,.+

  - name: SeverityField
    type: str
    description: |
      Name of the alert field that will receive the value from the matching
      `IncludeFilter` row's `Severity` column. Leave empty to omit severity
      from alerts entirely.
    default: Severity

imports:
  - Server.Monitor.Errors.Alert

sources:
  - query: |
      // Helper function for a nested array loop (just to rename "_value"):
      LET LabelSpecs(Specs) = SELECT _value AS Spec
        FROM foreach(row=Specs)

      // linter: invalid_arg:unlabelled|labelled
      LET MonitoredArtifacts = SELECT *
        FROM combine(unlabelled={
          SELECT NULL AS Label,
                 _value.artifact AS Artifact
          FROM foreach(row=get_client_monitoring().artifacts.specs)
        },
                     labelled={
          SELECT *
          FROM foreach(row=get_client_monitoring().label_events,
                       query={
          SELECT _value.label AS Label,
                 Spec.artifact AS Artifact
          FROM LabelSpecs(Specs=_value.artifacts.specs)
        })
        })
        WHERE IncludeArtifact(artifact=Artifact, filter=IncludeFilter)
      
      // NOTE: This is output once at the start of the query and is not going
      // to be up to date with the actual list of monitored artifacts, which
      // will be queried for every tick in clock() later:
      LET _ <= SELECT
          if(condition=Label,
             then=log(level='DEBUG',
                      message='Monitoring event artifact %v for label "%v"',
                      args=(Artifact, Label),
                      dedup=-1),
             else=log(level='DEBUG',
                      message='Monitoring event artifact %v',
                      args=Artifact,
                      dedup=-1))
        FROM MonitoredArtifacts

      LET LabelIsMonitored(Label, Labels) = if(
          condition=Label,
          then=Label IN Labels,
          else=true)

      // Poll logs for every currently active client monitoring artifact and emit
      // only entries that match filters:
      LET EventErrors(StartTime) = SELECT *
        FROM foreach(row={
          SELECT client_id AS ClientId,
                 labels
          FROM clients()
        },
                     query={
          SELECT *
          FROM foreach(row=MonitoredArtifacts,
                       query={
          SELECT Artifact,
                 Label,
                 ClientId,
                 client_time AS Timestamp,
                 level AS Level,
                 message AS Message
          FROM monitoring_logs(
            client_id=if(condition=LabelIsMonitored(Label=Label, Labels=labels),
                         then=ClientId),
            artifact=Artifact,
            start_time=StartTime)
          WHERE IncludeLog(artifact=Artifact,
                           level=level,
                           message=message,
                           filter=IncludeFilter)
           AND NOT IncludeLog(artifact=Artifact,
                              level=level,
                              message=message,
                              filter=ExcludeFilter)
        },
                       async=true,
                       workers=5)
        },
                     async=true,
                     workers=2)

      // Build the final alert context from the matching log row, including a
      // human-readable log timestamp:
      LET AlertArgs = RemoveEmptyStrs(Item=SeverityArgs + ErrorContext) +
          dict(
            dedup=DedupInterval,
            name=format(format='Client event query error in %v', args=Artifact),
            Level=Level,
            Error=Message,
            `Log timestamp`=TimestampString(Timestamp=timestamp(string=Timestamp)),
            Artifact=Artifact,
            ClientId=ClientId)

      // On each interval, inspect only the recent log window and raise alerts for
      // matching entries:
      SELECT *
      FROM foreach(row={
          SELECT Unix
          FROM clock(period=Period)
        },
                   query={
          SELECT *, alert(`**`=AlertArgs) AS Alert
          FROM EventErrors(StartTime=Unix - Period)
        })````