Create SMuRF Hammer Agent - #1092
Conversation
|
actually, if it's possible for an OCS panel to display data from an ocs feed, we wouldn't need another monitor process and could get this info from the det-controller agents (with a config change for the streamers). |
8b26187 to
47c0adb
Compare
|
@BrianJKoopman How do you usually proceed for testing new agents? Can I run this on a LAT smurf server in parallel to the production OCS and try hammering slots? |
mhasself
left a comment
There was a problem hiding this comment.
Non-exhaustive review -- looks good for the most part but:
- You'll need to add a page to the docs, especially including configuration information.
- Consider starting agent name with Smurf or Pysmurf so that it sorts better in agent lists (e.g. in ocs sidebar; hostmanager lists). E.g.
SmurfHammerAgent - It's nice to require privs=2, but functionally that will make this inert right now because we don't run access control at the site! Can be changed later.
I don't think we want to release this to the world without access control. What's the timescale for that? |
|
Thanks for taking a look! I renamed the agent and added to the docs, but holding off on changing the permissions until we decide on the path forward. If escalating permissions is not supported right now, I don't feel like it's very risky to deploy this without locking it down. I seems to me like anyone with access to ocs-web can already do a lot of damage (potentially a lot worse than hammering SMuRF). btw I'll be out on vacation next week, so if this converges on that timescale and someone wants to push it through, don't wait on me! |
Fair point @tristpinsm! |
|
In anticipation of needing to bring up all the smurf crates at site sometime in the next week, I'd like to be ready to deploy and test this agent when that happens. Here's how I think that would work: On a given smurf server:
to the list of agents.
Does that look correct? |
31c6ca6 to
6a7403a
Compare
Should probably use unique instance id per server, even at this stage. Btw, you don't have to edit the default.yaml in situ. In similar situations, I'll often duplicate the default.yaml to a scratch directory, set OCS_CONFIG_DIR to point to that scratch dir, and launch. So as not to mess with the stable configuration.
Only --instance-id arg would be needed. It should auto-detect the other things (assuming you've set OCS_CONFIG_DIR to something sensible for this user).
Yes, should work. And note you could run this on any machine as long as it knows how to find a default.yaml that connects to the right crossbar server -- doesn't have to be on the smurf server in question. Could be daq-lat. |
|
Some notes from testing today: I found some bugs in I had tried to set things up for the agent to work in a docker container, which required mounting the docker socket and adding the docker group so that jackhammer in the container can control containers on the host. This aspect seemed to be working, but I was running into a whack-a-mole of missing server configuration (missing random scripts in the path, ssh authentification to the shelf manager, ...) so I gave up on that for now. If we want it in a docker container, I think that would be doable, but need to spend more time identifying all the things it needs. Running as a process on the host, it worked fine, except the monitor process could not query the state of slots because epics is not installed on the host (I had debugged this yesterday in a docker container and the monitor is working). Another nice thing about the docker option is it was easy to pipe logs to Loki. Presumably you can do this for host agents as well, but I didn't know how. Options then for deploying this are:
I also wanted to test out the ocs-web panel for the agent. I built a docker image from my ocs-web branch and picked just the LAT crossbar entry out of the ocs-deployment-configs repo to use as a config file. I was able to start the instance of ocs-web on smurf-so8-lat, but in the web browser it didn't display any agents and complained that it was not connected to the crossbar. I checked that I could reach the crossbar address from within the container, but didn't investigate any further. |
The ocs-web docker doesn't need to contact crossbar -- it's actually just serving files (html/js/css) to your web browser. It is your web-browser that needs to be able to reach crossbar. So if that's running on the host system, then use the same addresses that you'd use for an agent on the host system. I usually develop ocs-web on a local machine, and ssh port forward the one OCS port that is needed to get the connection to crossbar. |
|
OK, with Matthew's guidance I was able to start a test ocs-web instance and validate that the panel can display the monitoring state of the slots. I didn't trigger a hammer yet via ocs-web because it wasn't a good time. I think the only remaining thing to sort out is whether to deploy this in a docker container or directly running on the host.
Option 2 is more work, but it may make it easier to maintain / deploy changes in the future. On the other hand, we probably want to continue supporting jackhammer on the host (unless we make it an alias that just drops you into a docker container). Any thoughts on how to deploy the agent? |
|
I vote we run it on host and install epics on all the smurvers. |
|
I had another opportunity to test now with the power briefly having gone down. This time with
I hammered using the Still need to figure out how to deploy the agent / send logs to Loki, but I think this is ready for a final review. |
|
bumping this for a review, would be great to get it out |
|
Sorry to be MIA on this one so far -- and thanks @mhasself and @msilvafe for the comments so far. I'll get you comments hopefully by end of day tomorrow, if not then certainly by the end of the week. A couple of quick comments:
|
|
Oh, and can you merge in the latest from |
New agent wrapping sodetlib's jackhammer hammer() function, exposing it as an OCS task with the same options (slots, no_reboot, no_dump, skip_setup, dump_rogue). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Store the traceback in session.data['error'] on failure so clients can inspect what went wrong. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Use the new hammer() return value to populate session.data with succeeded_slots and failed_slots. Distinguish full success, partial success, and total failure in the task return message. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
for more information, see https://pre-commit.ci
52b7dcf to
25df41a
Compare
|
done. Installing EPICS is a bit of a pain unfortunately. |
BrianJKoopman
left a comment
There was a problem hiding this comment.
Thanks for working on this, a few comments below.
- remove TimeoutLock - register feed with buffer time - partial success returns False - log exceptions, don't put in session data - tweak docs
for more information, see https://pre-commit.ci
52ca026 to
7eaf244
Compare
BrianJKoopman
left a comment
There was a problem hiding this comment.
This looks good, thanks for making those changes. One last question -- how involved was Claude Code in writing this? I see the co-commits by Claude here. Can you describe how it was used during this agent's development?
I'm trying to follow our new AI policies, which state:
Code and text generated by AI must be described as being so, with the tool used acknowledged. For example, in a pull request where AI assistance was used, you may write “AI Usage: The content of this pull request was generated by Claude Code”. It is helpful to also include information on exactly how the AI was used in this process, e.g., “The AI developed the code in process_fits.py, and I integrated it into main_system.py”.
You can find the full policy linked here: https://simonsobservatory.org/about/policies/
This is our first AI assisted code in the repo, so I might work out some other details related to this before I end up merging.
|
The commits labelled "Co-Authored-By: Claude" were generated using Claude code, and refined by me. Claude was used to build the initial structure of the agent and its documentation and to add features that were deemed necessary along the way. |
BrianJKoopman
left a comment
There was a problem hiding this comment.
Approving. Will merge when checks pass. I just added a note in the source code and a badge in the docs to note that this agent was developed with AI tools.
A more formal policy is to follow in a separate PR. The text/badge here might change when that happens, but this will do for now.
Thanks for the work on this!
Here is a first draft for a
jackhammeragent. It should allow individual slots to fail and track the results.Currently, I think you could set up an OCS-web panel that would display the result for each slot. This would only update to display the last jackhammer run however (and I guess only show the slots that were requested in that run).
If we wanted to have a continuously updating panel, we could add a monitor process to the agent that queries the
SmurfApplication.SystemConfiguredregister, modeled on the ACU agent. I'm not super familiar with OCS, so let me know if that is a good idea.Requires simonsobs/sodetlib#514.
AI Disclosure: