Alarms
Capabilities
Alarms capabilities
Capability |
Support |
Comment |
|---|---|---|
Create alarm from SNMP-TRAP |
SNMP-TRAP supports alarm creation through LUA rules. Refer to the dedicated chapter for more details. |
|
Create alarm from scripts |
Alarms can be created using scripts, allowing for the implementation of complex use cases. See Create alarms with dynamic data via scripting. |
|
Create alarm from callback (output) with static data |
Alarms can be triggered based on value changes via callbacks. See Create an alarm on value change (static data). |
|
Create alarm from callback (output) with dynamic data |
The context triggering the output is limited during alarm creation. However, this limitation can be bypassed using scripts. See Create alarms with dynamic data via scripting. |
|
Create alarm from IP-RCT |
Refer to AE IP-RCT for more information. |
|
Alarm deduplication |
Deduplication prevents duplicate alarms by treating them as a single occurrence. See Deduplication. |
|
Chat/journal linked to an alarm |
Each alarm includes a journal where users can add comments or log events. |
|
Add journal entries based on actions (e.g., escalation) |
Journals, as collections, support interactions such as automatic entry creation on specific actions. |
|
Define multiple severity levels |
Alarms include severity levels, each with a color code and specific behavior. See Severity. |
|
Modify alarm severity (escalation) from the GUI |
Severity can be manually changed from the GUI. Any associated action will be triggered upon severity change. |
|
Time-based severity escalation |
Severities can have escalation or expiration policies. For example, a minor alarm can escalate if not acknowledged within 30 minutes. See Expiration - clearing or escalating. |
|
Lock severity level for specific alarms |
Alarms can be locked to a specific severity, useful for managing noisy or redundant alarms. See Lock a severity level. |
|
Acknowledge alarms |
Alarms can be acknowledged to indicate they have been seen or handled. Acknowledgment is independent of severity levels. |
|
Mute or put an alarm in maintenance mode |
Alarms can be silenced for a defined period to avoid disturbances. Requires configuration. See Define alarms silencing rules with collections, Manage the maintenance state of alarms. |
|
Enrich alarms on creation (pre-insertion rules) |
Alarms can be enriched with additional information upon creation using pre-insertion rules. See Enrich alarms - pre-insertion rules. |
|
Periodic or manual evaluation of active alarms |
Active alarms can be re-evaluated periodically or manually, potentially modifying their state or triggering actions. See Evaluate alarms at regular intervals or by manual trigger. |
|
Automatic alarm removal after expiration |
Alarms can be automatically resolved when their expiration time is reached, based on severity settings. See Expiration - clearing or escalating. |
|
Update alarm content from scripts or GUI |
Alarms can be edited via scripts or GUI actions. Updates will trigger relevant callbacks. |
|
Add tags to alarms |
Alarms can be tagged via scripts or GUI actions. Updates will trigger relevant callbacks. |
|
Generate values based on filtered alarm pool |
You can generate values (e.g., count) based on filtered alarms (e.g., by severity). |
|
React to alarm creation/modification/deletion |
Alarms value can be used to react to change. See Extract data from alarms to a value. |
|
Generate values based on historical alarm pool |
Not currently supported; historical alarm pool is unavailable. |
|
Define parent-child relationships (grouping) |
BETA - See Grouping. Allows alarm grouping. For instance, alarms under a power outage can be suppressed when the parent alarm is active. |
|
Join/leave alarm groups from GUI or scripts |
BETA - See Grouping. Allows alarms to be grouped or ungrouped manually or programmatically. |
|
Handle value errors |
BETA - Allows creating alarms on value transitions to |
Frontend (table) capabilities
Generic features for tables
Capability |
Support |
Comment |
|---|---|---|
Customize menu |
||
Choice data to display (filter) |
Filter can be used to choose the data displayed to a user. See menus and toolbar |
|
Define a default filter by table |
A unique default filter can be set for each view. See Default filter |
|
Define a default filter by user or group of user |
The filter cannot be set by user or group of user, only by dashboard. But the local storage can be used to store the filter for a user. |
|
Set views or filter by URL parameters |
The filter can be set by URL parameters, which allows creating links from other widgets or dashboards. See URL parameters |
Concept
Alarms are a way to represent events that can be requested based on filters and views.
Note
The lifecycle of an alarm is defined between it’s creation and suppression. All operations done between these two instant is stored on one history entry. If the alarm is recreated later on a new lifecycle begins again. This means that a serial can have multiple entry on the history.
Alarms persistence
The current and past alarms are stored on a Mongo database.
Note
We provide, by default, a MongoDB instance with the alarms module. This is not recommended for production as the instance is configured for testing purpose without any access management.
Warning
MongoDB defines collections which are not the Collections defined by OnSphere. MongoDB collections are the equivalent of SQL tables.
Alarms are stored using five collections.
Buffer : This collection is used as a buffer for fast insertion. Elements in this collection are not yet considered to be part of the alarms.
Live : This collection stores the current alarms. All the actions possible on alarms are applied on this collection (escalate, add journal entry, …). It’s used to generate alarms lists views and logics.
History : This collection stores the whole history of received alarms. Elements in this collection are used to generate the alarms history lists views. The history only keeps the most recent alarms, the older ones are moved to the archives collections.
Archive: These collections store the older parts of the history. These collections are automatically generated for a defined time period (For example, archiveFrom01012023To03312023 for the alarms between the 01 january 2023 and 31 mars 2023).
Understand the alarms pipeline
Two ways are available to insert alarms in the database : standard insertion and direct insertion.
Direct insertion
Direct insertion is used for internal and scripted alarms insertion. This is used for inserting alarms directly into the database, without any pre-processing mechanisms. Alarms inserted this way are inserted one by one and the insertion result is returned to the caller, so this is not designed for efficiently inserting loads of alarms.
Standard insertion
The standard insertion mechanism, on the other hand, is designed to efficiently insert batches of alarms, and it allows pre-processing alarms right before they are inserted so they can be filtered out or enriched with additional information.
Standard insertion pipeline is decomposed into 3 steps :
Buffer
Incoming alarms are first inserted into a buffer collection without any pre-processing operation to keep those insertions as efficient as possible. This buffer is stored on the disk, so alarms are guaranteed to be handled by the system once they are inserted in it, even if the osp-alarms module restarts.
Every 500ms, or when the buffer has enough alarms in it, a fixed sized batch of alarms is taken from the buffer and passed into the rest of the pipeline.
Pre-insertion
The batch of alarms taken from the buffer is then given to the pre-insertion rules. Those pre-insertion rules evaluate a condition for each alarm of the batch to check if they must handle the alarm or not, and all pre-insertion rules matching an alarm will be applied to it, and so on for each alarm of the batch. Alarm not matching any pre-insertion rules are simply kept for insertion.
Pre-insertion rules are able to either
insert a new alarm (directly or injecting it into the buffer)
remove an alarm
update the alarm and forward it further in the pipeline
drop the alarm.
or any combination of those.
Operations execution
Once all the alarms have been handled by the pre-insertion rules, all operations requested by pre-insertion rules are applied to the database using a transaction. In particular, alarms that have not been filtered out are inserted into the live collection database.
Deduplication
An alarm is identified by a unique serial number, which serves as the sole criterion for detecting deduplication. The deduplication process ensures that when two alarms share the same serial, they are treated as a single alarm with updated values. The serial number is generated by the configuration of each alarm creation.
Default alarm fields
The content of an alarm as store on the database and displayed on the front-end.
Field name |
Description |
Type |
|---|---|---|
serial |
Unique identifier for an alarm (i.e., all alarms with same serial content will be merged). |
string |
firstTimestamp |
The first occurrence timestamp of the alarm as a long representing the nanoseconds. |
long |
lastTimestamp |
The last occurrence timestamp of the alarm as a long representing the nanoseconds. |
long |
severity |
The alarm severity (numeric value). As defined on the severity.ospp. |
integer |
summary |
A brief description of the alarms. |
string |
location |
The location of the alarm. For example: |
string |
source |
The origin of the alarm. For example: |
string |
tags |
An array of tags associated with the alarm. A tag is represented as a String. |
string[] |
additionalData |
Any other useful information as a key/value pair. |
map[string, object] |
acknowledged |
The acknowledge status (boolean value). |
boolean |
hideUntil |
A timestamp in nanoseconds to indicate the time until the alarm should be hidden. |
long |
forceHide |
A boolean to indicate if the alarm must be hidden. |
boolean |
journal |
A list of entries to store the actions done, comment on the current state and so on. |
JournalEntry[] |
count |
The total number of alarms using this serial. |
integer |
origSeverity |
The alarm original severity (numeric value) used to handle the locked severity. |
integer |
isSeverityLocked |
Flag indicating if the severity is locked. |
boolean |
highestSeverity |
The highest severity set on the alarm. |
integer |
lowestSeverity |
The lowest severity set on the alarm. |
integer |
parent |
The parent of the alarm or group. |
string |
children |
The list of the children of a group. This field is only available on a group. |
string[] |
History
The lifecycle of an alarm is defined between it’s creation and suppression. All operations done between these two instant is stored on one history entry. If the alarm is recreated later on a new lifecycle begins again. This means that a serial can have multiple entry on the history.
The history is a collection who contains all the actions done on a alarms.
When displayed on the front end, the alarms are rebuilt to match History alarm view. This is the same format as Default alarm fields plus some information like the which operation was done, when and an unique id per lifecycle.
Detailed content of an history
Field name |
Description |
|---|---|
_id |
Unique identifier of an history entry. |
serial |
Unique identifier of an alarm. |
creationTime |
The timestamp of the creation of the alarm as a long representing the nanoseconds. Define the start of the lifecycle. |
suppressionTime |
The timestamp of the suppression (archiving) of the alarm as a long representing the nanoseconds. Define the end of the lifecycle. It can be null when the alarm still exist on live. |
numberOfOccurrences |
The number of occurrence (creation of a new alarm with the same serial) during the lifecycle of the alarm. |
numberOfOperations |
The number of operations (action like create, acknowledge, …) during the lifecycle of the alarm. |
operations |
The list of the operations done on the alarm during it’s life cycle. |
Note
The operations are stored on the operations collection. The field operationId is used to match between the history and operations.
The following mongo pipeline allow to print the history with the operations.
db.history.aggregate([
{
"$lookup": {
"from": "operations",
"localField": "operations",
"foreignField": "operationId",
"as": "operations"
}
}
])
Operation detailed content
Field name |
Description |
|---|---|
operationTime |
The moment when the operation was done in nanoseconds. |
source |
The source of the operation. Who triggered it. |
data |
A optional data bag representing what was modified by the operation. |
The following operation types exist:
CREATE
ACKNOWLEDGE
UNACKNOWLEDGE
ESCALATE
LOCK_SEVERITY
UNLOCK_SEVERITY
TAG
UNTAG
ADD_JOURNAL_ENTRY
SET_SUMMARY
SET_LOCATION
SET_SOURCE
SET_ADDITIONAL_DATA
ADD_ADDITIONAL_DATA
REMOVE_ADDITIONAL_DATA
CLEAR_ADDITIONAL_DATA
SET_HIDE_UNTIL
SET_FORCE_HIDE
DELETE
JOIN_GROUP
LEAVE_GROUP
GENERATE_GROUP
REMOVE_GROUP
History alarm view
The view is rebuilt every 2 seconds by default for each changed alarms during this interval.
The interval can be configured on module.alarms with the field historyReconstructionInterval.
Field name |
Description |
Type |
|---|---|---|
serial |
Unique identifier for an alarm (i.e., all alarms with same serial content will be merged). |
string |
firstTimestamp |
The first occurrence timestamp of the alarm as a long representing the nanoseconds. |
long |
lastTimestamp |
The last occurrence timestamp of the alarm as a long representing the nanoseconds. |
long |
severity |
The alarm severity (numeric value). As defined on the severity.ospp. |
integer |
summary |
A brief description of the alarms. |
string |
location |
The location of the alarm. For example: |
string |
source |
The origin of the alarm. For example: |
string |
tags |
An array of tags associated with the alarm. A tag is represented as a String. |
string[] |
additionalData |
Any other useful information as a key/value pair. |
map[string, object] |
acknowledged |
The acknowledge status (boolean value). |
boolean |
hideUntil |
A timestamp in nanoseconds to indicate the time until the alarm should be hidden. |
long |
forceHide |
A boolean to indicate if the alarm must be hidden. |
boolean |
journal |
A list of entries to store the actions done, comment on the current state and so on. |
JournalEntry[] |
count |
The total number of alarms using this serial. |
integer |
origSeverity |
The alarm original severity (numeric value) used to handle the locked severity. |
integer |
isSeverityLocked |
Flag indicating if the severity is locked. |
boolean |
highestSeverity |
The highest severity set on the alarm. |
integer |
lowestSeverity |
The lowest severity set on the alarm. |
integer |
parent |
The parent of the alarm or group. |
string |
children |
The list of the children of a group. This field is only available on a group. |
string[] |
operationType |
The operation name that generated this alarm. |
string |
operationTime |
The timestamp of the operation. |
integer |
historyId |
The id of the corresponding history |
string |
Buffered alarms
The buffered alarms are waiting to be processed by the pre-insertion rules if any before joining live. Their format is different from a Alarm.
Field name |
Description |
Type |
|---|---|---|
serial |
Unique identifier for an alarm (i.e., all alarms with same serial content will be merged). |
string |
timestamp |
The occurrence timestamp of the alarm as a long representing the nanoseconds. |
long |
severity |
The alarm severity (numeric value). As defined on the severity.ospp. |
integer |
summary |
A brief description of the alarms. |
string |
location |
The location of the alarm. For example: |
string |
source |
The origin of the alarm. For example: |
string |
tags |
An array of tags associated with the alarm. A tag is represented as a String. |
string[] |
additionalData |
Any other useful information as a key/value pair. |
map[string, object] |
acknowledgement |
The acknowledge status (boolean value). |
boolean |
hideUntil |
A timestamp in nanoseconds to indicate the time until the alarm should be hidden. |
long |
forceHide |
A boolean to indicate if the alarm must be hidden. |
boolean |
initialJournal |
A list of entries to store the actions done, comment on the current state and so on. |
string[] |
parent |
The parent of the alarm or group. |
string |
Ignored alarms
The ingored alarms are the one rejected by the pre-insertion by using operations.ignore(). It use the same field as Buffered alarms.
Field name |
Description |
Type |
|---|---|---|
ignoreTimestamp |
The occurrence timestamp of the alarm as a date. |
IsoDate |
Note
This collection is not accessible from the front-end. It must be accessed via the mongo client directly.
Failure alarms
The failure alarms are the one that fail during the pre-insertion. It use the same field as Buffered alarms.
Field name |
Description |
Type |
|---|---|---|
failueTimestamp |
The occurrence timestamp of the alarm as a date. |
IsoDate |
Note
This collection is not accessible from the front-end. It must be accessed via the mongo client directly.
Create an alarm on value change (static data)
Overview
This system allows for generating static alarms based on value changes via a callback (trigger). The alarm is triggered by a Lua expression that evaluates to a boolean, with the context based on the value associated with the callback that triggered this output. If the condition is met, an action specified in the then section is executed. However, the content of the generated alarms is static and cannot be dynamically linked to the trigger’s value. This process relies on a fixed mapping between the trigger and the alarm actions, without the ability to dynamically customize the alarms based on specific values.
Use-case
Link the state of a monitoring device to an alarms to detect outage.
Create a basic scenario to generate an alarm for an outage with minimal data.
Usage
The if field must be a boolean evaluation with a LUA expression , if the field is evaluated to true the then part will be executed and will create an alarm. Otherwise the else condition will be executed.
If the condition encounters a non-existent case, no action will be taken. Note that the evaluation is static and has no knowledge of the context of the callback that is invoking the output.
The configuration of the alarms generation can use the content of the value to define some fields with a Value accessor.
The following fields support the extraction:
serial
summary
source
location
timestamp
The others are statically defined.
Example
Create alarms with dynamic data via scripting
Scripting allowing to create alarms see Run scripts from script
Severity
The severity represents the criticality of alarms. It is possible to define as many severities as needed but to ease integration, some severities are default provided.
Hint
When an unknown severity is used, the default one will be used to replace it.
Default severity
Name |
Value |
Default |
ItemId |
|---|---|---|---|
Clear |
100 |
root.alarms.severities.clear |
|
Intermediate |
200 |
root.alarms.severities.intermediate |
|
Warning |
300 |
root.alarms.severities.warning |
|
Minor |
400 |
root.alarms.severities.minor |
|
Major |
500 |
root.alarms.severities.major |
|
Critical |
600 |
root.alarms.severities.critical |
Customize Severity
Severity levels can be customized using three configuration files:
severity.web - Defines the display style for alarm tables.
severity.ospp - Specifies the severity levels.
severity.alarms - Links to the alarms module and sets the default severity.
Tip
When defining severities, keep in mind that the number of levels might increase over time. Using spaced numerical values (e.g., 100, 200, 300…) is a good practice, as it allows you to insert new levels more easily.
Tip
By default, OnSphere uses a low numerical value to indicate a less critical severity. This behavior can be inverted by changing the default sort order of severity.
Each severity level is associated with a numerical value, allowing for operations such as:
Checking if a severity is higher or lower than another.
Sorting alarms based on their severity.
Expiration - clearing or escalating
A severity can define a policy for an automatic escalation or expiration. For example, a clear alarm can be automatically removed after 2 hours or a minor alarm can be escalated to a major one if it was not acknowledged after 30 minutes.
The expiration field allows you to set a timer to automatically change a value (e.g., escalate or resolve an alarm).
For more details, see severity.alarms.
Lock a severity level
It is possible to lock an alarm severity. When the severity of an alarm is locked, it won’t change even if new alarms are received afterward with another severity.
Locking an alarm severity can be useful in two cases (among others):
when you want to change the severity of an alarm that is generated frequently (using alarm severity escalation will only change the severity until a new alarm is received)
when receiving “flipping alarms” (i.e., alarms that are often generated with a short period of time)
Note
Once the severity is unlocked, it will be updated with the last received severity (a flipping alarms switching from a clear severity to a critical severity, if locked to intermediate, once unlocked will be set back to either clear or critical depending on the last one received).
Enrich alarms - pre-insertion rules
Before insertion
During the insertion process, a series of scripts can be run to analyze the incoming alarms and trigger different mechanisms:
Add or modify alarms based on values, current alarms or state.
Reject alarms.
Pre-insertion rules
When inserting an alarm, a pre-insert rule can be used to apply conditional operations before the alarm is inserted.
It is possible to have multiple pre-insert rules in which case they are applied to alarms from highest to lowest priority (1 is higher than 10). They can be chained to from a full pre-processing pipeline :
A pre-insert rule is composed of a series of functions calls sequenced as follow :
Called functions are declared and defined in a lua script referenced in the pre-insert rule configuration file.
Warning
At least one pre-insertion rule must call operations.create() in order to create the alarm.
Functions signatures
The signatures of functions used by pre-insertion rules must be as follow:
Note
Below signatures descriptions are formatted as such functionName(param1Name: param1Type, param2Name: param2Type, ...).
The functionName define the configurationAttributeName in the file pre-insert.alarms.
When defining the function in the script LUA, you can name it freely (e.g shouldThisAlarmBeHidden, initialize, …).
In pre-insert.alarms, the configurationAttributeName must contains the name of the function define in the script LUA matching the signature.
- class pre-insertion()
- pre-insertion.beginBatch()
Optional function that allow to initialize the new pre-insertion pipeline.
- pre-insertion.for(alarm: BufferedAlarm)
Filter the alarm that will be processed by this rule.
- Arguments:
alarm – A buffered alarm aka the alarm that enter this rule and the result of the previous rules modifications if any. This is a static value.
- Returns:
A boolean to indicate if the rule must continue.
- pre-insertion.predicate(alarm: BufferedAlarm)
Generate a MongoDB filter that will be use to extract the FilterResult.
- Arguments:
alarm – A buffered alarm aka the alarm that enter this rule and the result of the previous rules modifications if any. This is a static value.
- Returns:
A string representing a MongoDB filter.
- pre-insertion.is(alarm: BufferedAlarm, filterResult: FilterResult)
Define the execution flow for the pre-insertion rule.
- Arguments:
alarm – A buffered alarm aka the alarm that enter this rule and the result of the previous rules modifications if any. This is a static value.
filterResult – FilterResult
- Returns:
A boolean to trigger
thenExecuteiftrueorelseExecuteiffalse.
- pre-insertion.thenExecute(alarm: BufferedAlarm, filterResult: FilterResult, operations: Operations)
The processing of current alarm to apply then the
isreturntrue.- Arguments:
alarm – A buffered alarm aka the alarm that enter this rule and the result of the previous rules modifications if any. This is a static value.
filterResult – FilterResult
operations – Operations that return a modifiable alarm for the chosen action. If the action is
forward, the modified alarm will be used to trigger the next rule.
- pre-insertion.elseExecute(alarm: BufferedAlarm, filterResult: FilterResult, operations: Operations)
The processing of current alarm to apply then the
isreturnfalse.- Arguments:
alarm – A buffered alarm aka the alarm that enter this rule and the result of the previous rules modifications if any. This is a static value.
filterResult – FilterResult
operations – Operations that return a modifiable alarm for the chosen action. If the action is
forward, the modified alarm will be used to trigger the next rule.
- pre-insertion.endBatch()
Optional function that allow to execute a function at the end of the pre-insertion pipeline.
Operations
Operations accessible to process the alarms on the pre-insertion pipeline:
Warning
As they are default provided, they MUST be called prefixed with “operations:”. i.e calling “create()” will not do what you want, you must call “operations:create()”.
- class operations()
- operations.create()
Trigger a Direct insertion at the end of the pre-insertion pipeline. The alarm generate by this call can be modified to change it’s content.
It’s used generally as the final step of the pre-insertion rules (lowest priority) to create the alarm.
- Returns:
The current editable buffered alarm.
- operations.insert()
Trigger a Standard insertion at the end of the pre-insertion pipeline. The alarm generate by this call can be modified to change it’s content.
- Returns:
The current editable buffered alarm.
- operations.forward()
Forward the current alarm to the next pre-insertion rule. The alarm generate by this call can be modified to change it’s content. It will be forwarded to the next rule with the modification done.
- Returns:
The current editable buffered alarm.
- operations.ignore()
Deprecated since version 1.4.0.
Ignore the current alarm to the next pre-insertion rule.
- operations.remove(criteria: string)
Deprecated since version 1.4.0.
Finds alarms matching a MongoDB criteria that will be removed from the live collection.
FilterResult
Some (c.f associated signatures section) of the functions you define and declare in alarms lua script take a FilterResult parameter. Its attributes are as follow:
Field |
Description |
Type |
|---|---|---|
count |
The number of alarms matching the filter |
Integer |
highestSeverity |
The highest severity of all matching alarms |
Integer |
lowestSeverity |
The lowest severity of all matching alarms |
Integer |
latestSeverity |
The latest severity of all matching alarms |
Integer |
firstTimestamp |
The first timestamp of all matching alarms |
Timestamp |
lastTimestamp |
The last timestamp of all matching alarms |
Timestamp |
hideUntil |
The last value of the field hideUntil |
Integer |
acknowledge |
The last value of the field acknowledge |
Boolean |
Evaluate alarms at regular intervals or by manual trigger
Action rules
When it is needed to run a task on alarms, either triggered manually of periodically, an action rule can be used. Action rules are composed of a series of functions calls sequenced as follow :
Called functions are declared and defined in a lua script referenced in the action rule configuration file.
Functions signatures
The signatures for the methods used by the actions rules must be as follow:
Note
Below signatures descriptions are formatted as such functionName(param1Name: param1Type, param2Name: param2Type, ...).
The functionName define the configurationAttributeName in the file action.alarms.
When defining the function in the script LUA, you can name it freely (e.g shouldThisAlarmBeHidden, initialize, …).
In action.alarms, the configurationAttributeName must contains the name of the function define in the script LUA matching the signature.
- class actionRule()
- actionRule.criteria(data: Map)
Generate a MongoDB filter that will be use to extract the FilterResult.
- Arguments:
data – User data provided by the trigger (front-end, script or timer).
nilwhen trigger is a timer.
- Returns:
A string representing a MongoDB filter.
- actionRule.beginBatch(data: Map)
Optional function that allow to initialize the new action rule.
- Arguments:
data – User data provided by the trigger (front-end, script or timer).
nilwhen trigger is a timer.
- actionRule.execute(alarm: Alarm, data: Map, actions: Actions)
The action to perform on each matching alarm extracted with the
criteria.
- actionRule.endBatch(data: Map)
Optional function that allow to execute a function at the end of the action rule.
- Arguments:
data – User data provided by the trigger (front-end, script or timer).
nilwhen trigger is a timer.
Actions
Actions are default provided and are made available (as parameters) from some (c.f associated signatures section) of the functions you define and declare in action rule’s lua script :
Warning
As they are default provided, they MUST be called prefixed with “actions:”. i.e calling “acknowledge()” will not do what you want, you must call “actions:acknowledge()”.
- class actions()
- actions.create(serial: string, severity: int)
Prepare an alarm for a Direct insertion at the end of the action rule. The alarm generate by this call can be modified to change it’s content.
- Arguments:
serial – The serial of alarm to create.
severity – The severity of alarm to create.
- Returns:
- actions.insert(serial: string, severity: int)
Prepare an alarm for a Standard insertion at the end of the action rule. The alarm generate by this call can be modified to change it’s content.
- Arguments:
serial – The serial of alarm to create.
severity – The severity of alarm to create.
- Returns:
- actions.remove()
Remove the alarm currently being processed.
- actions.acknowledge()
Acknowledge the alarm currently being processed.
- actions.unacknowledge()
Unacknowledge the alarm currently being processed.
- actions.lock_severity(severity: int)
Lock the severity of the alarm currently being processed.
- Arguments:
severity – The severity to which the alarm is locked.
- actions.unlock_severity()
Unlock the severity of the alarm currently being processed.
- actions.tag(tags: string[])
Tag the alarm currently being processed.
- Arguments:
tags – List of the tags to add.
- actions.untag(tags: string[])
Untag the alarm currently being processed.
- Arguments:
tags – List of the tags to remove.
- actions.escalate(severity: int)
Change the severity of the alarm currently being processed.
- Arguments:
severity – The severity to set.
- actions.edit(summary: string, location: string, source: string, hideUntil: int, forceHide: boolean, additionalData: object)
Edit the content of the alarm currently being processed.
Updates any of the fields summary, location, source, hideUntil, forceHide and additionalData of the alarm. Any field of additionalData can be added or updated, but they cannot be removed. The value null can be provided for fields that do not need to be updated and
{}for the additional data.- Arguments:
summary – The new summary.
location – The new location.
source – The new source.
hideUntil – Timestamp in ns.
forceHide – Status of forceHide.
additionalData – The new additionalData.
- actions.journal(serials: string[], message: string, user: string, status?: string)
Adds a journal entry to the alarm.
- Arguments:
serials – List of targeted alarms serials.
message – The message to add.
user – The name to use in the user field.
status – The status of the entry: success, info, warning or error. (default is info)
Some scripts can be triggered either periodically or manually. They can add, modify or remove alarms based on parameters, values, current alarms or state.
Extract data from alarms to a value
Context
A modification of one or multiple alarms can trigger the publication of one or multiple values. The publication can generate the following values :
The alarms current count
The minimum or maximum severity
The alarms count for a severity
The list of the modified or created alarms
Every values generated by the alarms are based on a filter.
Empty clear severity
If an alarm filter returns no results, we may still want to display a severity level (e.g., to indicate everything is fine). To handle this, filters can have an “empty clear severity,” which sets the severity for empty results.
Filter
A filter is used to observe a subset of alarms. For example, to react only for alarms located in Switzerland.
A filter defines a dynamic sub-set of alarms based on information available inside them. It can be used to only show alarms related to a location or type of device.
A filter can also be used to display a count, maximum severity, minimum severity or count for a specific severity. An overview of the system can be easily created with these information.
The Filter placeholder are available for creating a filter.
Note
With the Grouping, the group of an alarm matching the filter will be shown even if it doesn’t match it.
Query
Filters queries are created using mongoDB query operators.
The simplest query filter is the empty filter {}. This query will return the alarms without filtering.
A slightly more advanced filter could be to only retrieve alarms from a particular source:
{
"source": {
"$eq": "localhost"
}
}
This filter will return all alarms with the localhost source. It can be abbreviated as :
{
"source": "localhost"
}
Use-case
Count alarms above a severity.
Count alarms by a location.
Values
Values provided by alarms are linked to a filter.
ALARM_COUNT: The count of the alarm matching the filter.MAX_SEVERITY: The maximum severity of the alarms matching the filter.MIN_SEVERITY: The minimum severity of the alarms matching the filter.ON_ALARM: The list of the created, updated and deleted alarms since the last publication.SEVERITY_COUNT: The count of alarms with the defined severity.
Note
Computing filters values requires performing requests on the database.
Additionally, ON_ALARM values requires the module to keep a list of active alarms.
The memory footprint of the module will therefore increase for every ON_ALARM value that is configured.
In order to get the best performance, you should only define filter values if you need them.
Grouping
Warning
This feature is currently in beta. It may change in a future version without prior notice. See the Beta Features page for the full list of beta features and their planned release. If you’re using this feature, we encourage you to share your feedback to help with the evaluation process.
A group is a set of alarms (or groups) linked together with a parent-children relationship.
It has properties similar to alarms:
A serial
A summary
A Severity
A location
A source
An occurrence
Acknowledgement
Tags
A journal
…
Each of these properties can be defined by inheritance from children or statically (See Field generation).
As some actions available on alarms might not be relevant for a group with inherited field, the group can be configured to delegate or ignore actions (See Possible actions).
Link between alarms and groups
A group can have several children, including groups.
However, an alarm or group can only have one parent.
Field generation
Generation applies to the list of all children for extraction, addition, AND, OR…
This allows the content of the group’s fields to be dynamic and reflect the state of its members.
The most obvious use-case is for the group to inherit the highest severity of its children. It can also be used to give an overview of the summaries.
Following generators are available :
Generator name |
Description |
Remarks |
|---|---|---|
INITIAL |
Defines an initial value when the group is generated, which can then be modified |
Modifications are made using the same actions as for other alarms |
MAXIMUM |
Extracts the maximum value from a field |
|
MINIMUM |
Extract the minimum value of a field |
|
SUM |
Adds the contents of fields |
Only on numbers |
AND |
Logical AND between all the value from the children field |
Only on booleans |
OR |
Logical OR between all the value from the children field |
Only on booleans |
FROM_MAX_SEVERITY |
Extracts the value from the child field with the highest severity |
|
FROM_MIN_SEVERITY |
Extracts the value from the child field with the lowest severity |
|
FROM_LATEST |
Extract the value from the child field with the highest occurrence (lastTimestamp) |
|
FROM_OLDEST |
Extract the value from the child field with the lowest occurrence (firstTimestamp) |
|
CONCATENATE |
Concatenate the value from the child fields with or without a prefix |
Generate a string like “Prefix: value1, value2”. Can only be used to set a field containing a string |
MERGE_ARRAY |
Merge children’s arrays fields to create a union out of them. |
|
MERGE_OBJECT |
Merge the object of the children’s field |
Limitations will be applied depending on the type of alarm field. For example, a severity can only be an integer.
Possible actions
Actions available for alarms are also available for groups but depending on the configuration of the group some of them are meaningless. For example, escalate a group where the severity is defined by the maximum of its children.
To handle these cases, the configuration of a group allows to define the scope of actions.
Option name |
Description |
Remarks |
|---|---|---|
ON_NOBODY |
No action will be performed |
|
ON_CHILD |
Action only performed on children |
|
ON_HIMSELF |
Action only performed on the group |
Only if no inheritance in the field definition |
ON_BOTH |
Action performed on group and children |
Only if no inheritance in the field definition |
If the action on affects children and one of the children is a group, the group’s configuration will be taken into account.
Note
A group cannot be removed with the ARCHIVE operation. To remove a group (not defined statically on the configuration), the action REMOVE_GROUP must be used. When doing so, the children of the removed group will inherit their parent.
Note
If a field is not generated with INITIAL, the configuration dispatcher will refuse the configuration if an action is ON_HIMSELF or ON_BOTH for the same field.
Usage
Monitor an infrastructure
The state of an infrastructure can be represented on OnSphere. When this state is not correct an alarm can be generated to notify users that a problem appeared. It is also possible for an infrastructure to report problem with snmp-trap or webhook and create alarms.
Examples
Manage the maintenance
When a device has a problem, it can be useful to inhibit the alarms until the technician was able to fix the device.
Value error handler
The value error handler allow the module to generate alarms when any value (from all module) is publish with an error state.
It is disable by default and can be enable by defining valueErrorHandler on module.alarms.
When the value is not in error anymore, the alarm can be update by the following operation:
NOOP (default): No operation done
ACK: Acknowledge the corresponding alarm
ESCALATE: Escalate the corresponding alarm to the defined severity
DELETE: Remove the corresponding alarm
The configuration of the alarms generation can use the content of the value to define some fields with a Value accessor.
The following fields support the extraction:
serial
summary
source
location
timestamp
The others are statically defined.
Lua script function
Warning
Lua functions defined for alarms module behavior cannot be defined as “local”.
To ease integration, alarms Lua scripts provide the following functions and objects :
- class lua.log()
This object can be used to log messages.
Use the symbol
{}and a list of values to inject values inside the text message.log.trace("test: {}", {{name = "Automate 1"}})
Will give the message
test: Automate 1The current log level depends on the scriptLogLevel defined inside the module.alarms configuration file.
- log.trace(message: string, args: any[])
- Arguments:
message – String message with place holder
{}.args – The list of element to replace in the message.
- log.debug(message: string, args: any[])
- Arguments:
message – String message with place holder
{}.args – The list of element to replace in the message.
- log.info(message: string, args: any[])
- Arguments:
message – String message with place holder
{}.args – The list of element to replace in the message.
- log.warn(message: string, args: any[])
- Arguments:
message – String message with place holder
{}.args – The list of element to replace in the message.
- log.error(message: string, args: any[])
- Arguments:
message – String message with place holder
{}.args – The list of element to replace in the message.
- class lua.values()
This object is used to manipulate the timestamp of an alarm.
- values.get(id: string)
This function allows getting a value in the hierarchy. The value must be declared in the action or pre-insertion to be able to access it.
- Arguments:
id – The ItemId of the value
- Returns:
A value
- class lua.store()
Allow to store and retrieve shared user data into the database during the pre-insert and action rule script.
- store.get(identifier: string)
Retrieve the content store with the identifier.
- Arguments:
identifier – User defined string to identify the content.
- Returns:
The content stored or a
nilif nothing was found.
- store.set(identifier: string, content: any)
Store the content with the provided identifier.
- Arguments:
identifier – User defined string to identify the content.
content – Any data define by the user
- Returns:
The previous value or
nil.
- class lua.actionRules()
Execute a external script from Evaluate alarms at regular intervals or by manual trigger.
- class lua.script()
Execute a external script from Scripting.
- script.run(id: string, data?: Table, timeout?: int)
Example:
script.run('root.test.script') script.run('root.test.script', {1, 2, 3}) script.run('root.test.script', {}, 10)
- Arguments:
id – ItemId of the script to execute.
data – The argument to pass to the rule.
timeout – The maximum time to wait a response in milliseconds.
- Returns:
A table with the following content:
success: boolean to indicate if the execution was successful.
message: a message returned by the script.
content: a table containing the result of the script.
- script.runBlind(id: string, data?: Table, timeout?: int)
Example:
script.runBlind('root.test.script') script.runBlind('root.test.script', {1, 2, 3}) script.runBlind('root.test.script', {}, 10)
- Arguments:
id – ItemId of the script to execute.
data – The argument to pass to the rule.
timeout – The maximum time to wait a response in milliseconds.
- class lua.collections()
Makes a request to the Collections.
- collections.list(schemaId: string, pageSize: int, pageNumber: int)
- Arguments:
schemaId – The ItemId of the schema.
pageSize – The size of the page.
pageNumber – The page number.
- Returns:
A table with the following content:
totalCount: The number of result found.
collections: The list of each entry found limited by the pageSize.
- collections.listWithFilter(schemaId: string, pageSize: int, pageNumber: int, filterId: string)
- Arguments:
- Returns:
A table with the following content:
totalCount: The number of result found.
collections: The list of each entry found limited by the pageSize.
- collections.listWithCustomFilter(schemaId: string, pageSize: int, pageNumber: int, filter: string)
- Arguments:
schemaId – The ItemId of the schema.
pageSize – The size of the page.
pageNumber – The page number.
filter – The custom filter to use. Define as a mongoDB filter.
- Returns:
A table with the following content:
totalCount: The number of result found.
collections: The list of each entry found limited by the pageSize.
- collections.get(schemaId: string, documentId: string)
- Arguments:
schemaId – The ItemId of the schema.
documentId – The MongoDB identifier of the wanted document.
- Returns:
The document found.
- collections.getWithFilterId(schemaId: string, filterId: string)
- collections.getWithCustomFilter(schemaId: string, filter: string)
- Arguments:
schemaId – The ItemId of the schema.
filter – The custom filter to use. Define as a mongoDB filter.
- Returns:
The first document found.
- collections.insert(schemaId: string, data: Map<string, object>)
- Arguments:
schemaId – The ItemId of the schema.
data – The content of the document to insert.
- collections.update(schemaId: string, data: Map<string, object>)
If data doesn’t contain an _id, it will create a new element instead.
- Arguments:
schemaId – The ItemId of the schema.
data – The field to update on the document.
- collections.updateDiff(schemaId: string, documentId: string, data: List<Map<string, object>>)
- class lua.Timestamp()
This object is used to manipulate the timestamp of an alarm.
- Timestamp.isOlderThan(other: Timestamp)
The method can be used to determine if the current timestamp is older than the other one.
- Arguments:
other –
Timestamp()
- Timestamp.isNewerThan(other: Timestamp)
The method can be used to determine if the current timestamp is newer than the other one.
- Arguments:
other –
Timestamp()
- Timestamp.now()
Generate a
Timestamp()reprensting this moment.- Returns:
- Timestamp.from(timestamp: int)
Generate a
Timestamp()from epoch nanoseconds.- Arguments:
timestamp – A timestamp in nanosecond.
- Returns:
- Timestamp.getValue()
This method returns the underlying number of nano seconds since the Unix epoch.
- Returns:
A timestamp in nanoseconds.
- Timestamp.plusMillis(value: int)
- Arguments:
value – The number of milliseconds to add.
- Timestamp.plusSeconds(value: int)
- Arguments:
value – The number of seconds to add.
- Timestamp.plusMinutes(value: int)
- Arguments:
value – The number of minutes to add.
- Timestamp.plusHours(value: int)
- Arguments:
value – The number of hours to add.
- Timestamp.plusDays(value: int)
- Arguments:
value – The number of days to add.
- Timestamp.minusMillis(value: int)
- Arguments:
value – The number of milliseconds to substract.
- Timestamp.minusSeconds(value: int)
- Arguments:
value – The number of seconds to substract.
- Timestamp.minusMinutes(value: int)
- Arguments:
value – The number of minutes to substract.
- Timestamp.minusHours(value: int)
- Arguments:
value – The number of hours to substract.
- Timestamp.minusDays(value: int)
- Arguments:
value – The number of days to substract.
- Timestamp.fromat(format: string, epochNano: int, timezone: string)
This method returns a formatted date based on the timezone. For example, the call Timestamp.format(“yyyy-MM-dd hh:mm:ss”, 1644307200000000000, “CET”) will produce 2022-02-08 09:00:00. The timezone is UTC by default.
- Arguments:
epochNano – The timestamp extract with
getValue().timezone – The string representing the timezone.
format – The format string with the following symbol:
Symbol
Meaning
Presentation
Examples
G
era
text
AD; Anno Domini; A
u
year
year
2004; 04
y
year-of-era
year
2004; 04
D
day-of-year
number
189
M/L
month-of-year
number/text
7; 07; Jul; July; J
d
day-of-month
number
10
Q/q
quarter-of-year
number/text
3; 03; Q3; 3rd quarter
Y
week-based-year
year
1996; 96
w
week-of-week-based-year
number
27
W
week-of-month
number
4
E
day-of-week
text
Tue; Tuesday; T
e/c
localized day-of-week
number/text
2; 02; Tue; Tuesday; T
F
week-of-month
number
3
a
am-pm-of-day
text
PM
h
clock-hour-of-am-pm (1-12)
number
12
K
hour-of-am-pm (0-11)
number
0
k
clock-hour-of-am-pm (1-24)
number
0
H
hour-of-day (0-23)
number
0
m
minute-of-hour
number
30
s
second-of-minute
number
55
S
fraction-of-second
fraction
978
A
milli-of-day
number
1234
n
nano-of-second
number
987654321
N
nano-of-day
number
1234000000