osp-alarms

Performance consideration

Following points can impact the performance of the alarms handling.

  • The size of alarms.

  • The complexity of filters.

  • The number of occurrences.

The principal symptom will be a long time to display a change on the front-end. The grafana dashboard can help detecting these points.

Migration

The migration must be done to adapt the alarm to the new format.

Database migration from the previous version must be done for:

  • 1.x.y to 2.0.0

The alarms module includes an auto migration feature but it can be slow if the database has a lot of data. For big database (more than 3 GB of data), we strongly recommend using the migration tool.

Warning

The alarms module must not be started in the new version before the migration tool is started. This can cause errors during the migration process.

The tool is available as a docker image nexus.onsphere.ch/osp-alarms-migration:<version>. It has the following parameters:

usage: Alarms migration [-h] [--connectionString CONNECTIONSTRING]
                        [--database DATABASE]

Migrate alarms and history from 1.x to 2.0.0.
This tool will replace live with the content of deduplicated.
Then it will convert each alarms on the buffer, history and archives.


named arguments:
  -h, --help             show this help message and exit
  --connectionString CONNECTIONSTRING
                        The  connection   string   for   mongo   database.
                        (default: mongodb://localhost)
  --database DATABASE    Name of the database (default: alarms)

Warning

No automatic rollback is available.

Note

You can see collections created on the alarms database by running show collections. Collections to migrate are the following :

  • archiveFrom*To*

  • buffer

  • failure

  • history

  • ignored

  • live

  • deduplicated (removed)

Following are possible:

  • Runs on the command line docker run --rm -it nexus.onsphere.ch/osp-alarms-migration:<version> --connectionString mongodb://modules_mongodb_osp-mongo-1/?replicaSet=sdn0 --database=alarms history live

    Note

    This only works if MongoDB is accessible from outside the stack. Using a network defined inside a Docker stack from a container outside the stack is not allowed.

  • Runs as a one time service (from inside the stack or as an external service with the network as osp-stack-1_mongo):

    osp-alarms-migration:
      image: nexus.onsphere.ch/osp-alarms-migration:<version>
      deploy:
        replicas: 1
        restart_policy:
          condition: none
      command: --connectionString mongodb://modules_mongodb_osp-mongo-1/?replicaSet=sdn0 --database=alarms history live
      networks:
        - mongo
    

You can use the following strategy to do the migration:

  • Migrate all collection with the migration utility.

  • Migrate all collection other than live and buffer with the migration utility.

  • Don’t use the migration utility and let the module do all the work (ok for small installation). The module can restart multiple time during the migration.

List of configuration files

Filename

Short description

Format

Documentation

module.service

Each service is defined in a separate file and then combined into a unified compose stack file

yml

See the Swarm administration or Official documentation

module.alarms

The module description

json

module.alarms

severity.ospp

Define the property of a severity (name, description).

json

severity.ospp

severity.alarms

Define a severity for the alarms to use

json

severity.alarms

filter.alarms

Define the filter to apply on MongoDB.

json

filter.alarms

filter.ospp

Define the property of a filter (name, description).

json

filter.ospp

owner.alarms

Define a value generated based on the current alarms.

json

owner.alarms

view.alarms

Define the mapping of the fields between an alarms and the name define on the view.ospp.

json

view.alarms

view.ospp

Define the property of a view (name, description, fields)

json

view.ospp

pre-insert.alarms

Define a processing to execute when new alarms is inserted.

json

pre-insert.alarms

action.alarms

Define a processing to execute on a subset of the alarms either triggered periodically or manually

json

action.alarms

output.alarms

Generate an alarm when a value change.

json

output.alarms

Supervision Prometheus (Beta)

Execution Time

Metric

Labels

Type

Description

pre_insert_rule_execution_seconds

ItemId

histogram

Pre-insert rule execution time in seconds

archive_execution_seconds

None

histogram

Archive execution time in seconds

history_execution_seconds

None

histogram

History execution time in seconds

alarm_action_execution_seconds

None

histogram

Alarm action execution time in seconds

insertion_execution_seconds

None

histogram

Time to insert into the buffer

processing_execution_seconds

step

histogram

Processing time per step in seconds

Buffer & Batch Metrics

Metric

Labels

Type

Description

insertion_count_total

None

counter

Total alarms inserted into the buffer

buffer_size

None

gauge

Buffer size before processing (insertion)

processing_batch_size

  • batch

  • total_batch

gauge

Size of processing batch (insertion)

Processing Execution Created Timestamps

Metric

Labels

Type

Description

archive_execution_seconds_created

None

gauge

Timestamp of archive execution

history_execution_seconds_created

None

gauge

Timestamp of history execution

insertion_count_created

None

gauge

Timestamp of insertion count

insertion_execution_seconds_created

None

gauge

Timestamp of insertion execution

processing_execution_seconds_created

  • buffer

  • total_batch
    • batch

    • pre-alarms

    • alarm

      • check

      • aggregate

    • post-alarms

    • write

gauge

Timestamp of processing execution per step

Metric

Labels

Type

Description

action_rule_execution_seconds

itemId

histogram

Duration in seconds of the execution of each action rule, labeled by itemId.

alarm_action_execution_seconds

Metric

Labels

Type

Description

alarm_action_execution_seconds

acknowledge, escalate, lock, unlock, clear, tag, untag, journal, edit, create, insert

histogram

Duration in seconds of each alarm action execution, labeled by the corresponding action.

Front view

Metric

Labels

Type

Description

live_change_execution_seconds

step=find

histogram

Duration in seconds of the find operation on the database.

live_change_execution_seconds

step=compute

histogram

Duration in seconds to compute the difference between two alarm states, publishing only the changes to the front-end.

live_change_execution_seconds

filter

histogram

Label representing each applied filter.

An example for a Grafana dashboard : Dashboard

Environment Variables

All modules env variables

Variable Name

Default Value

Usage

PROMETHEUS_PORT

9100

The internal port used for Openmetrics exposition.

DISPATCHER_HOST

modules_configuration-dispatcher_main

The hostname on which the modules can fetch the configuration.

DISPATCHER_PORT

10000

The port used by the module to listen for configuration request.

USE_LEGACY_RESTART

Not set

When this flag is enabled, the container will terminate itself whenever a restart is needed (for example, after a configuration change). This is the legacy restart (before version 2.1.0). Otherwise, the system will simply restart the process running inside the container.

CUSTOM_JVM_OPTIONS

Not set

Allow to inject JVM options. See Configuration changes for more information.

ALLOW_VERSION_MISMATCH

Not set

Allow to disable the version check when a new configuration is published to a module. By default, a module will not be restarted if its configuration does not match its version.

Dedicated variables

None