Known bug

Service update of osp-web stuck with too much containers (3/1)

Description

This bug might appear during a configuration update. The module osp-web use service update strategy start-first which create the next container before killing the old one. Sometimes, the update algorithm may fail for some unknown reason and “pause” the service update. This state will remain until a new configuration is applied.

Symptom

Multiple module web are running concurrently.

$ docker service ls

uvmmlss17uen   osp-stack-1_modules_web_web-1                       replicated   3/1        osp-web:1.4.0

Workaround

Manually trigger a service update using portainer or command line:

docker service update --force osp-stack-1_modules_web_web-1

It is also possible to change update fail strategy setting in file modules/web/web-1/module.service

...
deploy:
    update_config:
    parallelism: 1
    delay: 0s
    order: start-first
+   failure_action: continue

Alarm module fail to build the history view

Description

The History alarm view process the history to prepare the view displayed in the alarms history table. This process can fail due to the size of the generated view when a alarm has a lot of operations in a history entry.

The problem occurs when the size of the rebuild view exceed 100MB.

Symptom

The history table will not show the last operation for an alarm after the interval define on the configuration.

A message similar to the one below will appear on the log of the module alarms.

com.mongodb.MongoCommandException: Command execution failed on MongoDB server with error 146 (ExceededMemoryLimit): 'PlanExecutor error during aggregation :: caused by :: $reduce would use too much memory and cannot spill' on server mongodb-0.mongodb-service:27017. The full response is {"ok": 0.0, "errmsg": "PlanExecutor error during aggregation :: caused by :: $reduce would use too much memory and cannot spill", "code": 146, "codeName": "ExceededMemoryLimit", "$clusterTime": {"clusterTime": {"$timestamp": {"t": 1774960232, "i": 1}}, "signature": {"hash": {"$binary": {"base64": "zuECBcEY12Fkzu8QuEMSbkgY8Jo=", "subType": "00"}}, "keyId": 7573668209333633025}}, "operationTime": {"$timestamp": {"t": 1774960232, "i": 1}}}

Diagnostic

To know if an alarm can cause this problem, you can compute the estimated size of the alarm view by multiplying the size of the alarm (in live) with the number of operation in the history entry.

If the multiplication of db.live.stats().avgObjSize and db.history.aggregate([{$project:{"opeCount":{$cond:{if:{$isArray:"$operations"},then:{$size:"$operations"},else:0}}}},{$sort:{"opeCount":-1}}]) is greater than 100MB. The alarm reconstruction will fail.

Note

Both request must be run on the mongodb shell.

Workaround

There is no workaround available yet.