osp-alarms
Performance consideration
Following points can impact the performance of the alarms handling.
The size of alarms.
The complexity of filters.
The number of occurrences.
The principal symptom will be a long time to display a change on the front-end. The grafana dashboard can help detecting these points.
Migration
The migration must be done to adapt the alarm to the new format.
Database migration from the previous version must be done for:
1.x.y to 2.0.0
The alarms module includes an auto migration feature but it can be slow if the database has a lot of data. For big database (more than 3 GB of data), we strongly recommend using the migration tool.
Warning
The alarms module must not be started in the new version before the migration tool is started. This can cause errors during the migration process.
The tool is available as a docker image nexus.onsphere.ch/osp-alarms-migration:<version>.
It has the following parameters:
usage: Alarms migration [-h] [--connectionString CONNECTIONSTRING]
[--database DATABASE]
Migrate alarms and history from 1.x to 2.0.0.
This tool will replace live with the content of deduplicated.
Then it will convert each alarms on the buffer, history and archives.
named arguments:
-h, --help show this help message and exit
--connectionString CONNECTIONSTRING
The connection string for mongo database.
(default: mongodb://localhost)
--database DATABASE Name of the database (default: alarms)
Warning
No automatic rollback is available.
Note
You can see collections created on the alarms database by running show collections. Collections to migrate are the following :
archiveFrom*To*
buffer
failure
history
ignored
live
deduplicated (removed)
Following are possible:
Runs on the command line
docker run --rm -it nexus.onsphere.ch/osp-alarms-migration:<version> --connectionString mongodb://modules_mongodb_osp-mongo-1/?replicaSet=sdn0 --database=alarms history liveNote
This only works if MongoDB is accessible from outside the stack. Using a network defined inside a Docker stack from a container outside the stack is not allowed.
Runs as a one time service (from inside the stack or as an external service with the network as osp-stack-1_mongo):
osp-alarms-migration: image: nexus.onsphere.ch/osp-alarms-migration:<version> deploy: replicas: 1 restart_policy: condition: none command: --connectionString mongodb://modules_mongodb_osp-mongo-1/?replicaSet=sdn0 --database=alarms history live networks: - mongo
You can use the following strategy to do the migration:
Migrate all collection with the migration utility.
Migrate all collection other than live and buffer with the migration utility.
Don’t use the migration utility and let the module do all the work (ok for small installation). The module can restart multiple time during the migration.
List of configuration files
Filename |
Short description |
Format |
Documentation |
|---|---|---|---|
module.service |
Each service is defined in a separate file and then combined into a unified compose stack file |
yml |
See the Swarm administration or Official documentation |
module.alarms |
The module description |
json |
|
severity.ospp |
Define the property of a severity (name, description). |
json |
|
severity.alarms |
Define a severity for the alarms to use |
json |
|
filter.alarms |
Define the filter to apply on MongoDB. |
json |
|
filter.ospp |
Define the property of a filter (name, description). |
json |
|
owner.alarms |
Define a value generated based on the current alarms. |
json |
|
view.alarms |
Define the mapping of the fields between an alarms and the name define on the view.ospp. |
json |
|
view.ospp |
Define the property of a view (name, description, fields) |
json |
|
pre-insert.alarms |
Define a processing to execute when new alarms is inserted. |
json |
|
action.alarms |
Define a processing to execute on a subset of the alarms either triggered periodically or manually |
json |
|
output.alarms |
Generate an alarm when a value change. |
json |
Supervision Prometheus (Beta)
Execution Time
Metric |
Labels |
Type |
Description |
|---|---|---|---|
pre_insert_rule_execution_seconds |
ItemId |
histogram |
Pre-insert rule execution time in seconds |
archive_execution_seconds |
None |
histogram |
Archive execution time in seconds |
history_execution_seconds |
None |
histogram |
History execution time in seconds |
alarm_action_execution_seconds |
None |
histogram |
Alarm action execution time in seconds |
insertion_execution_seconds |
None |
histogram |
Time to insert into the buffer |
processing_execution_seconds |
step |
histogram |
Processing time per step in seconds |
Buffer & Batch Metrics
Metric |
Labels |
Type |
Description |
|---|---|---|---|
insertion_count_total |
None |
counter |
Total alarms inserted into the buffer |
buffer_size |
None |
gauge |
Buffer size before processing (insertion) |
processing_batch_size |
|
gauge |
Size of processing batch (insertion) |
Processing Execution Created Timestamps
Metric |
Labels |
Type |
Description |
|---|---|---|---|
archive_execution_seconds_created |
None |
gauge |
Timestamp of archive execution |
history_execution_seconds_created |
None |
gauge |
Timestamp of history execution |
insertion_count_created |
None |
gauge |
Timestamp of insertion count |
insertion_execution_seconds_created |
None |
gauge |
Timestamp of insertion execution |
processing_execution_seconds_created |
|
gauge |
Timestamp of processing execution per step |
Metric |
Labels |
Type |
Description |
|---|---|---|---|
action_rule_execution_seconds |
itemId |
histogram |
Duration in seconds of the execution of each action rule, labeled by itemId. |
alarm_action_execution_seconds
Metric |
Labels |
Type |
Description |
|---|---|---|---|
alarm_action_execution_seconds |
acknowledge, escalate, lock, unlock, clear, tag, untag, journal, edit, create, insert |
histogram |
Duration in seconds of each alarm action execution, labeled by the corresponding action. |
Front view
Metric |
Labels |
Type |
Description |
|---|---|---|---|
live_change_execution_seconds |
step=find |
histogram |
Duration in seconds of the find operation on the database. |
live_change_execution_seconds |
step=compute |
histogram |
Duration in seconds to compute the difference between two alarm states, publishing only the changes to the front-end. |
live_change_execution_seconds |
filter |
histogram |
Label representing each applied filter. |
An example for a Grafana dashboard : Dashboard
Environment Variables
All modules env variables
Variable Name |
Default Value |
Usage |
|---|---|---|
PROMETHEUS_PORT |
9100 |
The internal port used for Openmetrics exposition. |
DISPATCHER_HOST |
modules_configuration-dispatcher_main |
The hostname on which the modules can fetch the configuration. |
DISPATCHER_PORT |
10000 |
The port used by the module to listen for configuration request. |
USE_LEGACY_RESTART |
Not set |
When this flag is enabled, the container will terminate itself whenever a restart is needed (for example, after a configuration change). This is the legacy restart (before version 2.1.0). Otherwise, the system will simply restart the process running inside the container. |
CUSTOM_JVM_OPTIONS |
Not set |
Allow to inject JVM options. See Configuration changes for more information. |
ALLOW_VERSION_MISMATCH |
Not set |
Allow to disable the version check when a new configuration is published to a module. By default, a module will not be restarted if its configuration does not match its version. |
Dedicated variables
None