Commit Graph
96 Commits
Author SHA1 Message Date
Nate Coraor a2d1625780 Merge remote-tracking branch 'upstream/release_21.01' into dev 2021-03-01 16:52:02 -05:00
davelopez 385611d3d7 Add unit test for fake job name guessing 2021-02-26 11:48:04 +01:00
guerler 427594be71 Fix test cases, naming issues 2021-02-19 14:56:49 -05:00
John Chilton 15bde671e7 Fix certain classes of workflow extraction on copied objects. 2021-01-15 23:59:06 -05:00
John Chilton ce9240e921 Rework refactor responses for structured messages... 2021-01-03 19:52:28 -05:00
John Chilton 6be4e2bdbd Rework refactoring API for better separation of concerns.
Thinner controller, pydantic model to separate options available for manager (which has to render a response) from those of the refactoring executor.
2021-01-03 19:52:28 -05:00
John Chilton 38fbd87c88 Use existing workflow_manager to load subworkflows in wf modules. 2021-01-03 19:52:28 -05:00
Bjoern Gruening 2cd2777746 adopt to code review 2021-01-02 18:48:32 +00:00
Bjoern Gruening dfba28a443 fix mutable arguments in tests 2021-01-02 18:48:31 +00:00
John Chilton fc2ce45767 API for structured workflow refactoring. 2020-12-29 22:00:32 -05:00
Nicola Soranzo 1ec1f515b3 Merge branch 'release_20.09' into dev 2020-12-09 19:35:39 +00:00
mvdbeek cdad7c79a9 Pass optional subworkflow inputs to editor
Fixes https://github.com/galaxyproject/galaxy/issues/10864
2020-12-08 12:29:59 +01:00
mvdbeek 2f5742c7c0 Merge branch 'release_20.09' into dev 2020-11-03 16:04:22 +01:00
mvdbeek 14ced3cbae Set FK on left side of one to many DCE/DC relation
I think this may fix a circular dependency between dataset_collection
and dataset_collection_element. This doesn't appear to be a problem
if dataset_collection has an id already, but the set of optimizations
that went into 20.09 may get us into the situation where that is not the
case.

I hope this fixes:
```
galaxy.job_execution.output_collect ERROR 2020-11-01 14:58:59,036 Problem gathering output collection.
Traceback (most recent call last):
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/engine/base.py", line 1284, in _execute_context
cursor, statement, parameters, context
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/engine/default.py", line 590, in do_execute
cursor.execute(statement, parameters)
psycopg2.errors.NotNullViolation: null value in column "dataset_collection_id" violates not-null constraint
DETAIL:  Failing row contains (13330967, null, 30637828, null, null, 0, ERR4597396__single).
The above exception was the direct cause of the following exception:
Traceback (most recent call last):
File "/opt/galaxy/server/lib/galaxy/job_execution/output_collect.py", line 156, in collect_dynamic_outputs
final_job_state=job_context.final_job_state,
File "/opt/galaxy/server/lib/galaxy/model/store/discover.py", line 286, in populate_collection_elements
self.flush()
File "/opt/galaxy/server/lib/galaxy/job_execution/output_collect.py", line 214, in flush
self.sa_session.flush()
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/orm/scoping.py", line 163, in do
return getattr(self.registry(), name)(*args, **kwargs)
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/orm/session.py", line 2523, in flush
self._flush(objects)
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/orm/session.py", line 2664, in _flush
transaction.rollback(_capture_exception=True)
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/util/langhelpers.py", line 69, in __exit__
exc_value, with_traceback=exc_tb,
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/util/compat.py", line 178, in raise_
raise exception
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/orm/session.py", line 2624, in _flush
flush_context.execute()
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/orm/unitofwork.py", line 422, in execute
rec.execute(self)
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/orm/unitofwork.py", line 589, in execute
uow,
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/orm/persistence.py", line 236, in save_obj
update,
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/orm/persistence.py", line 995, in _emit_update_statements
statement, multiparams
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/engine/base.py", line 1020, in execute
return meth(self, multiparams, params)
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/sql/elements.py", line 298, in _execute_on_connection
return connection._execute_clauseelement(self, multiparams, params)
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/engine/base.py", line 1139, in _execute_clauseelement
distilled_params,
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/engine/base.py", line 1324, in _execute_context
e, statement, parameters, cursor, context
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/engine/base.py", line 1518, in _handle_dbapi_exception
sqlalchemy_exception, with_traceback=exc_info[2], from_=e
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/util/compat.py", line 178, in raise_
raise exception
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/engine/base.py", line 1284, in _execute_context
cursor, statement, parameters, context
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/engine/default.py", line 590, in do_execute
cursor.execute(statement, parameters)
sqlalchemy.exc.IntegrityError: (psycopg2.errors.NotNullViolation) null value in column "dataset_collection_id" violates not-null constraint
DETAIL:  Failing row contains (13330967, null, 30637828, null, null, 0, ERR4597396__single).
[SQL: UPDATE dataset_collection_element SET dataset_collection_id=%(dataset_collection_id)s WHERE dataset_collection_element.id = %(dataset_collection_element_id)s]
[parameters: {'dataset_collection_id': None, 'dataset_collection_element_id': 13330967}]
(Background on this error at: http://sqlalche.me/e/gkpj)
```
2020-11-02 10:16:17 +01:00
Nicola Soranzo 9d74bba7fb Drop support for retired Python 3.5
Upgrade syntax using `pyupgrade --py36-plus` .

Manually drop several `six` imports.

Also:
- Remove broken pr_cache in scripts/bootstrap_history.py
- Fix broken prefix removal in lib/galaxy/tool_util/deps/mulled/mulled_build.py
2020-10-07 11:52:13 +01:00
John Chilton fe003d139e GA4GH search API. 2020-08-28 20:29:20 -04:00
John Chilton ed5bc4d1d6 API + UI for importing workflows from a GA4GH TRS server. 2020-08-21 12:06:54 -04:00
guerler 0547ec6cfc Do not automatically overwrite the output labels for subworkflows 2020-06-26 15:29:46 -04:00
mvdbeek 785cfce3fa Fix unit tests 2020-04-25 20:40:26 +02:00
John Chilton 26d5150245 Eliminate extensions=['input_collection'].
The existing hack of extensions=['input'] is bad enough and should work fine for collections. Not sure what past-John was thinking by adding this.
2019-12-11 12:26:32 -05:00
John Chilton 20d737cd09 Fix unit tests. 2019-12-10 21:21:31 -05:00
John Chilton fa637518e8 Implement yet another syntax for Galaxy Markdown.
The community took a vote (https://github.com/galaxyproject/galaxy/pull/8511#issuecomment-525886525) and the overwhelming favorite was this programming style declaration of a vaguely C-ish language embedded in quote fenced blocks.
2019-08-30 09:46:14 -04:00
John Chilton 664ab71653 Rework toward more RMarkdown-y syntax for Galaxy Markdown. 2019-08-30 09:43:15 -04:00
John Chilton ad650897d6 Implement workflow invocation reports.
Implement markdown backend and frontend components as well generator plugin framework to allow customizable workflow invocation reports.

Architectural Choices on the Client

Old-style embedded objects vs components

In theory, this could make really nice use of excellent VueJS components for dataset display, dataset collection display, workflow display, etc.. These aren't available yet, we don't even really have Backbone stuff that exist very well outside history panels, so this re-uses "components" from Galaxy Pages for datasets and reuses the collection display for history panels as a stand-alone display for collections (I got this trick from DIsplayStructured.vue). This isn't ideal and I know that but at least reusing things this way will provide a path forward for migrating both pages and this new Markdown language (which could easily replace or co-exist with pages HTML someday) together seamlessly as the real modern components become available and doesn't increase the overall work needed to integrate newer style components. In fact, this might even be the impetus for creating and polishing more of these components.

markdown-it vs markdown-it-vue

I tried this with markdown-it-vue also but it added very little in terms of reducing client code, obscured entirely how to attach plugins to the markdown rendering process, and brought in many, many more extra packages that we don't need or want.

Given there is no representation of invocations in the GUI and not editor for reports config this is a bit challenging still. But here is goes:

- Source Galaxy's virtualenv.
- Login to a user and grab an API key, set in the following command:
- ``GALAXY_TEST_EXTERNAL=http://localhost:8080 GALAXY_TEST_USER_API_KEY=38175dfc69d9009992b69efcca8b209b pytest test/api/test_workflows.py::WorkflowsApiTestCase::test_workflow_invocation_report_custom``
- In the web browser go to http://localhost:8080/api/invocations and grab the latest invocation ID. Replace it in the follow query string.
- Navigate to http://localhost:8080/workflows/invocations/report?id=c887f1d0da42bdfe
2019-08-30 09:43:01 -04:00
Nicola Soranzo e007cc928f Fix `items() and keys()` not supporting indexing on Py3 2019-08-18 19:25:44 +01:00
Martin Cech e65604bc06 Merge pull request #7556 from jmchilton/expression_tools
Implement expression tools and non-data tool outputs.
2019-03-20 18:52:59 -04:00
Dannon 65ddaefefa Merge pull request #7560 from jmchilton/refactor_security_helper
Refactor galaxy.web.security into galaxy.security.idencoding.
2019-03-20 11:18:28 -04:00
mvdbeek 5e8cbcd5b4 Allow connecting ExpressionTool outputs to inputs of correct type
This means we can now connect these tools in the workflow editor.
2019-03-20 12:55:04 +01:00
John Chilton 528497f7c4 Refactor galaxy.web.security into galaxy.security.idencoding.
I want to be able to use this class in galaxy.model and it doesn't have any dependnecies on "web" stuff so I'd like to move it out of galaxy.model and into galaxy.security. galaxy.model shouldn't have dependencies on galaxy.web (and doesn't after the recent 976f5ad367), so this pre-emptively ensures it doesn't require these dependencies for https://github.com/galaxyproject/galaxy/pull/7367.

This has the benefit of also eliminating any galaxy.web dependencies from the generic test base file (test/base/testcase.py) which has long been a goal of mine as well.

There are places we legimately use ID encoding and decoding for security (i.e. job files API) but for the most part it doesn't provide security for API endpoints and this causes confusion repeatedly, so I've started the process of renaming the contained class SecurityHelper to something I feel makes it clearer this is just about encoding and decoding IDs.
2019-03-19 14:28:31 -04:00
John Chilton 4a308130b3 Undo DynamicTool.tool_hash that is the part that is really half baked. 2019-03-19 14:22:41 -04:00
mvdbeek d36d1a25e6 Simple filter_output
This obviously doesn't work (yet) with filters that work
on the input dataset, but this already removes a lot of
outputs that arent' going to be produced based on static
options. Since exceptions default to priduing the dataset
this isn't much harm.
2019-03-05 11:26:08 +01:00
mvdbeek 667f8aa761 Only add collection_type if collection_type has been specified
We don't force the collection_type to be set in any way.
We might want to think about defaulting to list if not specified,
but that's not for now.
2019-01-28 17:12:50 +01:00
mvdbeek 6a4307609d Unit test resolve_collection_type 2019-01-23 15:25:44 +01:00
John Chilton 67efd072ae Track workflow step input definitions in our model.
We don't track workflow step inputs in any formal way in our model currently. This has resulted in some current hacks and prevents future enhancements. This commit splits WorkflowStepConnection into two models WorkflowStepInput and WorkflowStepConnection - normalizing the previous table workflow_step_connection on input step and input name.

In terms of current hacks forced on it by restricting all of tool state to be confined to a big JSON blob in the database - we have problems distinguishing keys and values when walking tool state. As we store more and more JSON blobs inside of the giant tool state blob - the worse this problem gets. Take for instance checking for runtime parameters or the rules parameter values - these both use JSON blobs that aren't simple values, so it is hard to tell looking at the tool state blob in the database or the workflow export to tell what is a key or what is a value. Tracking state as normalized inputs with default values and explicit attributes runtime values should allow much more percise state definition and construction.

This variant of the models would also potentially allow defining runtime values with non-tool default values (so default values defined for the workflow but still explicitly settable at runtime). The combinations of overriding defaults and defining runtime values were not representable before.

In terms of future enhancements, there is a lot we cannot track with the current models - such as map/reduce options for collection operations (https://github.com/galaxyproject/galaxy/issues/4623#issuecomment-389544980). This should enable a lot of that. Obviously there are a lot of attributes defined here that are not yet utilized, but I'm using most (all?) of them downstream in the CWL branch. I'd rather populate this table fully realized and fill in the implementation around it as work continues to stream in from the CWL branch - to keep things simple and avoid extra database migrations. But I understand if this feels like speculative complexity we want to avoid despite the implementation being readily available for inspection downstream.
2018-11-15 21:34:02 +01:00
John Chilton b1fe6fee93 Refactor collection mapping workflows toward independence from tools.
Refactor get_data_inputs into get_all_inputs and use the resulting dictionaries to reason about if collection mapping should occur during invocation of tools. Using these dictionaries instead of explicit tool input objects should allow reuse within other module types since they may produce the same interface.
2018-10-24 09:34:40 -04:00
Nicola Soranzo 6b3f672770 Replace deprecated assertEquals() method
Fix warnings during py34-unit tests like:

```
/home/travis/build/galaxyproject/galaxy/test/unit/workflows/test_extract_summary.py:46: DeprecationWarning: Please use assertEqual instead.
  self.assertEquals(job_dict[hda.job], [('out1', derived_hda_2)])
```

See e.g. https://travis-ci.org/galaxyproject/galaxy/jobs/438829277
2018-10-09 12:37:15 +01:00
mvdbeek 60c4162ffa Fix unit tests that access output label 2018-09-26 14:29:22 +02:00
mvdbeek a44baa5642 Look up collection_type if collection_type_source is set
We do this in the wf-editor and in the backend, so this works for
subworkflows. This should make quite some tools more usable in a
workflow scenario. Should fix
- https://github.com/galaxyproject/galaxy/issues/6514
- https://github.com/galaxyproject/galaxy/issues/6012
- https://github.com/galaxyproject/galaxy/issues/6569
- https://github.com/galaxyproject/galaxy/issues/1889
and perhaps more issues related to this.
2018-09-18 17:11:40 +02:00
John Chilton 7806a84d1c Implement runtime parameters in nested workflows. 2018-09-06 14:43:22 -04:00
mvdbeek 57f5b42974 py3: fix workflow extract summary 2018-07-05 02:49:51 +01:00
Nic Herndon 03a4172680 Fixed bugs with unit test on Python 3 2018-06-30 23:06:42 +00:00
guerler 028d07323d Adjust test case 2018-03-31 03:08:25 -04:00
John Chilton d5f98f4904 Formalize workflow invocation and invocation step outputs.
Workflow Invocations
--------------------

The workflow invocation outputs half of this is relatively straight forward. It is modelled somewhat on job outputs, output datasets and output dataset collections are now tracked for each workflow invocation and exposed via the workflow invocation API. This required adding new tables (linked to WorkflowInvocations and WorkflowOutputs) that track these output associations.

Previously one could imagine backtracking this information for simple tool steps via the WorkflowInvocationStep -> Job table, but for steps that have many jobs (i.e. mapping over a collection) or for non-tool steps such information was more difficult to recover (and simply couldn't be recovered from the API at all or even internally without significant knowledge of the underlying workflow).

Workflow Invocation Steps
-------------------------

Tracking the outputs of WorkflowInvocationSteps was not previously done at all, one would have to follow the Job table as well. A signficant downside to this is that one cannot map over empty collections in a workflow - since no such job would exist. Tracking job outputs for WorkflowInvocationSteps is not a simple matter of just attaching outputs to an existing table because we had no concept of a workflow step tracked - since there could be many WorklfowInvocationSteps corresponding to the same combination of WorkflowInvocation and WorkflowStep. That should feel wrong and that is because it is - when collections were added the possiblity of having many jobs for the same combination of WorkflowInvocation and WorkflowStep was added. I should have split WorkflowInvocationSteps into WorkflowInvocationSteps and WorkflowInvocationStepJobAssociations at that time but didn't. This commit now does it - effectively normalizing the ``workflow_invocation_step`` table by introducing the new ``workflow_invocation_step_job_association`` table.

Splitting up the WorkflowInvocationStep table this way allows recovering the mapped over output (e.g. the implicitly created collection from all the jobs) as well the outputs from the individual jobs (by walking WorkflowInvocationStep -> WorkflowInvocationStepJobAssociation -> Job -> JobToOutput*Association).

This split up involves failrly substantial changes to the workflow module interface. Any place a list of WorkflowInvocationSteps was assumed, I reworked it to just expect a single WorkflowInvocationStep. I vastly simplified recover_mapping to just use the persisted outputs (this was needed in order to also implment empty collection mapping in workflows). This also fixes a bug (or implements a missing feature) where Subworkflow moudles had no recover_mapping methods - so for instance if a tool that produces dynamic collections appeared anywhere in a workflow after a subworkflow step - that workflow would not complete scheduling properly.

Now that we have a way to reference the set of jobs corresponding to a workflow step within an invocation, we can start to track partial scheduling of such steps. This is outlined in https://github.com/galaxyproject/galaxy/issues/3883 and refactoring toward this goal is included here - including adding a state to WorkflowInvocationStep so Galaxy can determine if it has started scheduling this step and an index when scheduling jobs so it can tell how far into a scheduling things have gone as well as augmenting the tool executor to take a maximum number of jobs to execute and allow recovery of existing jobs for collection building purposes.

*Applications*

These changes will enable:

- A simple, consistent API for finding workflow outputs that can be consumed by Planemo for testing workflows.
- Mapping over empty collections in workflows.
- Re-scheduling workflow invocations that include subworkflow steps.
- Partial scheduling within steps requiring a large number of jobs when scheduling workflow invocations.
2017-11-30 10:02:01 -05:00
Nicola Soranzo 2cd95c48f6 Fix import order everywhere
- Add flake8-import-order to flake8 Pipfile and remove py27-lint-imports
  and py27-lint-imports-include-list tox envs
- Fix most E201 and E202 errors reported by flake8-import-order v0.15,
  but pin flake8-import-order to v0.14.3 until
  https://github.com/PyCQA/flake8-import-order/issues/123
  is fixed

This let us drop 2 jobs on Travis per each job.
2017-11-14 19:42:39 +00:00
E Rasche 32b85ecb59 Only permit yaml.safe_loading of data
Event trusted data, belt + suspenders method.
2017-09-25 11:39:12 +02:00
mvdbeek 64f8e74ade Fix subworkflow tests 2017-08-20 12:49:43 +02:00
Nicola Soranzo 21b44bf348 Fix all E201 and E202 style errors
using the following command:
```
autopep8 -i -r --exclude $(sed -e 's|^|./|' -e 's|/$||' .ci/flake8_blacklist.txt | paste -sd,) --select E201,E202 .
```
2017-08-17 11:35:39 +01:00
guerler 05687b8a19 Strip debug output 2017-02-22 13:03:43 -05:00
guerler ae751adac8 Adjust unit test to use labels instead of name through state 2017-02-22 12:14:23 -05:00
guerler 48ba525eb5 Merge branch 'dev' into fix_workflow_issues 2017-01-08 12:36:53 -05:00