Flush the job and its output datasets all at the same time in one transaction.
Inuitively it makes sense that this should work and the timings seem to improve, triming another quater or third of the submission time per job.
Before:
- galaxy.tools.execute DEBUG 2015-12-01 12:58:51,154 Tool [cat1] created job [21] (261.932 ms)
- galaxy.tools.execute DEBUG 2015-12-01 12:59:11,354 Tool [cat1] created job [22] (278.699 ms)
- galaxy.tools.execute DEBUG 2015-12-01 12:59:43,105 Tool [cat1] created job [23] (296.438 ms)
After:
- galaxy.tools.execute DEBUG 2015-12-01 13:22:00,649 Tool [cat1] created job [39] (265.547 ms)
- galaxy.tools.execute DEBUG 2015-12-01 13:22:27,720 Tool [cat1] created job [40] (214.905 ms)
- galaxy.tools.execute DEBUG 2015-12-01 13:21:40,225 Tool [cat1] created job [38] (198.936 ms)
- galaxy.tools.execute DEBUG 2015-12-01 13:22:44,076 Tool [cat1] created job [41] (213.096 ms)
... during tool execution. Build method once instead of in each function call, use imap instead of map since we don't need a list, remove some duplicated checks. Frankly this is all stuff Python is probably doing anyway - but in case it doesn't and just so the eye doesn't jump to these optimizations again.
Timings before and after for a section of tool action execute that includes this additon show that this might have a small effect.
Before:
- galaxy.tools.actions INFO 2015-12-01 13:11:39,446 Add outputs to history (127.379 ms)
- galaxy.tools.actions INFO 2015-12-01 13:12:06,598 Add outputs to history (137.029 ms)
- galaxy.tools.actions INFO 2015-12-01 13:12:23,931 Add outputs to history (118.489 ms)
After:
- galaxy.tools.actions INFO 2015-12-01 13:13:09,999 Add outputs to history (99.456 ms)
- galaxy.tools.actions INFO 2015-12-01 13:13:38,573 Add outputs to history (126.131 ms)
- galaxy.tools.actions INFO 2015-12-01 13:13:54,538 Add outputs to history (137.643 ms)
- galaxy.tools.actions INFO 2015-12-01 13:14:12,516 Add outputs to history (101.451 ms)
Just keep the dataset in the correct NEW state until it has actually been queued. Addresses FIXME comment added by James 7 years ago in https://github.com/galaxyproject/galaxy/commit/4c3db1af95fb0520960046ae549462aebd78b326.
This behavior feels correct to me, but it does have ramifications in the GUI. I had previously never actually seen a dataset in the "NEW" state.
Saves an extra flush per dataset, on sqlite this translates to 50ms per dataset for me.
- If a step has a label, display it in the step editor side panel title.
- Allow clicking the title to change the label.
- Display the label (if set) as the workflow node box title.
- Add icon to workflow node since the title might not be related to type anymore.
- Enforce unique labels accross the workflow in the editor.
- Qunit test cases for some of this behavior and other recent changes.
- Reduce duplication between initializing generic modules and tool modules.
- Switch tool_id to content_id as variable names throughout the client.
- Rename get_tool_id to get_content_id on workflow modules.
- Add some minimal documentation to the workflow module about get_content_id.
Downstream in the subworkflow commit I switch the over-the-wire communication to use content_id instead of tool_id also and use content_ids to refer to workflow ids in subworkflow moduls.