Enable flag: validation.run_core_services · Play tag: -t core_services
Checks that the OpenNebula control plane on the primary frontend is healthy and resilient: the XML-RPC API answers, every configured systemd unit (oned, gate, flow, Fireedge by default) is enabled at boot and survives a restart, the OS package manager has an OpenNebula repository configured, one-gate/one-flow/Fireedge listen on their configured ports, every registered KVM host survives a disable/enable monitoring cycle, and datastores and marketplaces are registered with sane capacity. All checks run with failed_when: false — a failing check is recorded in the report and the run continues.
The role runs only on the first host of the frontend group (the primary FE); it is skipped entirely, with a Core services validation = skipped — cluster discovery failed row, when the shared cluster discovery failed.
- Optional induced-failure hook fires if
odv_fail_in=core_services(orvalidation.debug.fail_in), then facts andservice_factsare gathered and thejc/jqpackages are installed via the shared tracked installer (manifested, so the sweep can remove them later). oneuser listverifies the XML-RPC API.- For each entry of
validation.core_services.service_list: record whether the unit is enabled at boot, restart it, and record whether it is running afterwards. Units not present on the FE get askipped - service not presentrow instead. oned -vrecords the version; the OS package sources are grepped for the configured OpenNebula repository (Debian and RedHat families).- On OpenNebula < 6.99 with
opennebula-scheduler.servicepresent, the legacy scheduler is stop/started andsched.logis checked foroned successfully contacted. - one-gate and one-flow listening ports are verified against their server config files;
opennebula-flowis additionally stop/start cycled. - If
validation.core_services.check_fireedge_uiis true and the unit exists, Fireedge is stop/start cycled, checked for boot enablement, its listening socket is polled (5 × 2 s), and the listen address is curl-ed. Otherwise all four Fireedge keys are recorded as skipped. - Every host registered in
onehost listis manifested (host_state), then cycled:onehost disable→enable→forceupdate→ poll up to 300 s (30 × 10 s) for the host to return to stateon. - Datastore and marketplace pools are listed; non-
systemdatastores must not report negativeFREE_MB. - An aggregate
Core services validationrow is computed:failedif any of this role's own report keys recordedfailed, elseok.
Ordering note (HA/RAFT): step 3 restarts opennebula.service on the primary FE. In an HA zone that is usually the RAFT leader, so the restart triggers a leader failover and destabilizes monitoring and the flow server for minutes. This is why the OneFlow smoke test is deliberately ordered before this role in playbooks/validation.yml, and why validation.one_flow.state_retries/state_delay may need raising on HA frontends when tests are re-run out of order.
- A deployed OpenNebula frontend reachable as the first host of the
frontend_groupinventory group, with the CLI tools (oneuser,onehost,onedatastore,onemarket) authenticated for the connecting user (host cycling runs asoneadminvia become). - systemd (the role relies on
service_factsandsystemctl), plusssandcurlon the FE. - Package-manager access to install
jcandjq(or have them preinstalled). - Registered hypervisor hosts whose monitoring can bring them back to state
onwithin 300 s of a disable/enable cycle. Do not run this while a host is fenced or down — the cycle check will fail on it. - Tolerance for brief service interruption: every listed unit is restarted, so the API/GUI blink, and on HA the RAFT leader fails over.
| Option | Default | Description |
|---|---|---|
validation.run_core_services |
true |
Enable/disable the whole test. |
validation.core_services.service_list |
opennebula.service, opennebula-gate.service, opennebula-flow.service, opennebula-fireedge.service (each as {name, desc}) |
Units to check for boot enablement and restart-resilience. desc is used to build the report keys. The reference inventory ships the same list; add entries to check more units. opennebula-scheduler.service needs no entry — it is checked automatically on < 6.99. |
validation.core_services.check_fireedge_ui |
true |
Run the four Fireedge GUI checks (stop/start, boot enablement, listening port, HTTP connect). When false or the unit is absent, they are recorded as skipped. |
validation.debug.fail_in |
unset | Set to core_services to trigger the induced-failure hook (harness testing). Equivalent per-run extra var: -e odv_fail_in=core_services. |
frontend_group |
frontend |
Inventory group whose first host runs the checks (global suite variable, not under validation.). |
Minimal inventory fragment (group_vars/all.yml):
validation:
run_core_services: true
core_services:
check_fireedge_ui: true
service_list:
- name: opennebula.service
desc: OpenNebula core (oned)
- name: opennebula-gate.service
desc: OpenNebula gate
- name: opennebula-flow.service
desc: OpenNebula flow
- name: opennebula-fireedge.service
desc: OpenNebula Fireedge GUIFull suite:
Copy
playbooks/report_input_vars.yml.exampletoplaybooks/report_input_vars.ymland fill in the report details first (see the main README).
make validation I=inventory/<env>/hosts.yaml ANSIBLE_ARGS='-e @playbooks/report_input_vars.yml'Only this test (plus the always-on report/sweep plays):
make validation I=inventory/<env>/hosts.yaml TAGS=core_services
# or directly:
hatch env run -e validation-default -- ansible-playbook -i inventory/<env>/hosts.yaml \
-t core_services -e @playbooks/report_input_vars.yml playbooks/validation.ymlReport keys recorded via verification_result (comments carry the error detail on failure):
XML-RPC API status—ok/failed.<desc> is enabled,<desc> restarted— one pair perservice_listentry present on the FE (ok/failed);<desc> not installed=skipped - service not presentfor absent units.Oned version— theoned -vbanner (rendered INFO) orfailed.Configured OpenNebula repository— the repo URL orfailed.OpenNebula scheduler can contact oned—ok/failed/skipped; recorded only on < 6.99.OpenNebula one-gate service listens on the server—ok/failed.Verify one-flow stops and starts,Verify one-flow service port is open on the server—ok/failed/skipped.Verify fireedge GUI stops and starts,Verify fireedge GUI is enabled to start automatically,Verify fireedge GUI listens port,Verify connection to a fireedge server—ok/failed, orskipped - fireedge UI check disabled or service not present; the connect check alone can also recordskipped - fireedge port not listeningwhen the port poll failed.Status of the registered KVM hosts— host id/name/state list,no hosts registered, orfailed.Put hosts offline and then set online. Hosts status—okonly if every host returned toonwithin 300 s;skipped - no hosts registeredotherwise applicable.Datastores list,Datastore capacity more than 0,Registered marketplaces— pool listings (INFO) andok/failedchecks. An empty marketplace pool isfailed.Core services validation— the aggregate row:failed(comment lists the failing keys) if any key above failed,okotherwise, orskipped — cluster discovery failed.
Badges in the HTML/PDF report follow the shared mapping: values containing failed/error render FAIL, skipped … disabled renders INFO, other skipped values render WARN, ok renders PASS, and free-text listings (version, pools, host tables) render INFO.
- Manifest: each registered host is manifested as type
host_statebefore cycling;jc/jqare manifested as typepackageonly if this run actually installed them. - always block (unconditional restore): every registered host is re-enabled — but only when its state is genuinely
dsbl/off/offline/disabled, followed byonehost forceupdate. A blindonehost enableon an already-enabled host would reset its monitoring to INIT and leave the cluster momentarily unschedulable for the next test, so healthy hosts are left untouched. Allservice_listunits, the legacy scheduler (< 6.99) and Fireedge are then (re)started regardless of check outcomes. - Leftovers: any host the always block failed to re-enable is recorded via
record_leftover(typehost_state) and rendered as a red LEFTOVER row in the report's Execution & Cleanup Summary.host_stateentries are deliberately excluded from the final sweep plan — restoration is this role's own job; the sweep only reports them. Manifested packages installed by this run are uninstalled by the final sweep. - rescue: catches harness bugs only (plumbing failures — individual checks cannot throw) and records
Core services validation = failedwith the failing task's detail. The play never aborts.