JtODAssoc mock is required because otherwise on construction the real
JtODAssoc will try to use its internal SQLAlchemy instrumentation to
setup the backref relationship to HDA, which does not exist in the mock;
hense the JtODAssoc itself needs to be a mock.
This fixes the original commit cceb5d2 (reverted in 7fa3313).
Remove unnecessary call to class_mapper(HDA)
Address SAWarning:
> relationship 'HistoryDatasetAssociation.creating_job_associations' will copy
> column history_dataset_association.id to column
> job_to_output_dataset.dataset_id, which conflicts with relationship(s):
> 'JobToOutputDatasetAssociation.dataset' (copies history_dataset_association.id
> to job_to_output_dataset.dataset_id). If this is not the intention, consider
> if these relationships should be linked with back_populates, or if
> viewonly=True should be applied to one or more if they are read-only. For the
> less common case that foreign key constraints are partially overlapping, the
> orm.foreign() annotation can be used to isolate the columns that should be
> written towards. The 'overlaps' parameter may be used to remove this
> warning. (Background on this error at: http
I think this may fix a circular dependency between dataset_collection
and dataset_collection_element. This doesn't appear to be a problem
if dataset_collection has an id already, but the set of optimizations
that went into 20.09 may get us into the situation where that is not the
case.
I hope this fixes:
```
galaxy.job_execution.output_collect ERROR 2020-11-01 14:58:59,036 Problem gathering output collection.
Traceback (most recent call last):
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/engine/base.py", line 1284, in _execute_context
cursor, statement, parameters, context
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/engine/default.py", line 590, in do_execute
cursor.execute(statement, parameters)
psycopg2.errors.NotNullViolation: null value in column "dataset_collection_id" violates not-null constraint
DETAIL: Failing row contains (13330967, null, 30637828, null, null, 0, ERR4597396__single).
The above exception was the direct cause of the following exception:
Traceback (most recent call last):
File "/opt/galaxy/server/lib/galaxy/job_execution/output_collect.py", line 156, in collect_dynamic_outputs
final_job_state=job_context.final_job_state,
File "/opt/galaxy/server/lib/galaxy/model/store/discover.py", line 286, in populate_collection_elements
self.flush()
File "/opt/galaxy/server/lib/galaxy/job_execution/output_collect.py", line 214, in flush
self.sa_session.flush()
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/orm/scoping.py", line 163, in do
return getattr(self.registry(), name)(*args, **kwargs)
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/orm/session.py", line 2523, in flush
self._flush(objects)
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/orm/session.py", line 2664, in _flush
transaction.rollback(_capture_exception=True)
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/util/langhelpers.py", line 69, in __exit__
exc_value, with_traceback=exc_tb,
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/util/compat.py", line 178, in raise_
raise exception
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/orm/session.py", line 2624, in _flush
flush_context.execute()
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/orm/unitofwork.py", line 422, in execute
rec.execute(self)
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/orm/unitofwork.py", line 589, in execute
uow,
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/orm/persistence.py", line 236, in save_obj
update,
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/orm/persistence.py", line 995, in _emit_update_statements
statement, multiparams
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/engine/base.py", line 1020, in execute
return meth(self, multiparams, params)
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/sql/elements.py", line 298, in _execute_on_connection
return connection._execute_clauseelement(self, multiparams, params)
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/engine/base.py", line 1139, in _execute_clauseelement
distilled_params,
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/engine/base.py", line 1324, in _execute_context
e, statement, parameters, cursor, context
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/engine/base.py", line 1518, in _handle_dbapi_exception
sqlalchemy_exception, with_traceback=exc_info[2], from_=e
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/util/compat.py", line 178, in raise_
raise exception
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/engine/base.py", line 1284, in _execute_context
cursor, statement, parameters, context
File "/opt/galaxy/venv/lib64/python3.6/site-packages/sqlalchemy/engine/default.py", line 590, in do_execute
cursor.execute(statement, parameters)
sqlalchemy.exc.IntegrityError: (psycopg2.errors.NotNullViolation) null value in column "dataset_collection_id" violates not-null constraint
DETAIL: Failing row contains (13330967, null, 30637828, null, null, 0, ERR4597396__single).
[SQL: UPDATE dataset_collection_element SET dataset_collection_id=%(dataset_collection_id)s WHERE dataset_collection_element.id = %(dataset_collection_element_id)s]
[parameters: {'dataset_collection_id': None, 'dataset_collection_element_id': 13330967}]
(Background on this error at: http://sqlalche.me/e/gkpj)
```
Upgrade syntax using `pyupgrade --py36-plus` .
Manually drop several `six` imports.
Also:
- Remove broken pr_cache in scripts/bootstrap_history.py
- Fix broken prefix removal in lib/galaxy/tool_util/deps/mulled/mulled_build.py
Implement markdown backend and frontend components as well generator plugin framework to allow customizable workflow invocation reports.
Architectural Choices on the Client
Old-style embedded objects vs components
In theory, this could make really nice use of excellent VueJS components for dataset display, dataset collection display, workflow display, etc.. These aren't available yet, we don't even really have Backbone stuff that exist very well outside history panels, so this re-uses "components" from Galaxy Pages for datasets and reuses the collection display for history panels as a stand-alone display for collections (I got this trick from DIsplayStructured.vue). This isn't ideal and I know that but at least reusing things this way will provide a path forward for migrating both pages and this new Markdown language (which could easily replace or co-exist with pages HTML someday) together seamlessly as the real modern components become available and doesn't increase the overall work needed to integrate newer style components. In fact, this might even be the impetus for creating and polishing more of these components.
markdown-it vs markdown-it-vue
I tried this with markdown-it-vue also but it added very little in terms of reducing client code, obscured entirely how to attach plugins to the markdown rendering process, and brought in many, many more extra packages that we don't need or want.
Given there is no representation of invocations in the GUI and not editor for reports config this is a bit challenging still. But here is goes:
- Source Galaxy's virtualenv.
- Login to a user and grab an API key, set in the following command:
- ``GALAXY_TEST_EXTERNAL=http://localhost:8080 GALAXY_TEST_USER_API_KEY=38175dfc69d9009992b69efcca8b209b pytest test/api/test_workflows.py::WorkflowsApiTestCase::test_workflow_invocation_report_custom``
- In the web browser go to http://localhost:8080/api/invocations and grab the latest invocation ID. Replace it in the follow query string.
- Navigate to http://localhost:8080/workflows/invocations/report?id=c887f1d0da42bdfe
I want to be able to use this class in galaxy.model and it doesn't have any dependnecies on "web" stuff so I'd like to move it out of galaxy.model and into galaxy.security. galaxy.model shouldn't have dependencies on galaxy.web (and doesn't after the recent 976f5ad367), so this pre-emptively ensures it doesn't require these dependencies for https://github.com/galaxyproject/galaxy/pull/7367.
This has the benefit of also eliminating any galaxy.web dependencies from the generic test base file (test/base/testcase.py) which has long been a goal of mine as well.
There are places we legimately use ID encoding and decoding for security (i.e. job files API) but for the most part it doesn't provide security for API endpoints and this causes confusion repeatedly, so I've started the process of renaming the contained class SecurityHelper to something I feel makes it clearer this is just about encoding and decoding IDs.
This obviously doesn't work (yet) with filters that work
on the input dataset, but this already removes a lot of
outputs that arent' going to be produced based on static
options. Since exceptions default to priduing the dataset
this isn't much harm.
We don't track workflow step inputs in any formal way in our model currently. This has resulted in some current hacks and prevents future enhancements. This commit splits WorkflowStepConnection into two models WorkflowStepInput and WorkflowStepConnection - normalizing the previous table workflow_step_connection on input step and input name.
In terms of current hacks forced on it by restricting all of tool state to be confined to a big JSON blob in the database - we have problems distinguishing keys and values when walking tool state. As we store more and more JSON blobs inside of the giant tool state blob - the worse this problem gets. Take for instance checking for runtime parameters or the rules parameter values - these both use JSON blobs that aren't simple values, so it is hard to tell looking at the tool state blob in the database or the workflow export to tell what is a key or what is a value. Tracking state as normalized inputs with default values and explicit attributes runtime values should allow much more percise state definition and construction.
This variant of the models would also potentially allow defining runtime values with non-tool default values (so default values defined for the workflow but still explicitly settable at runtime). The combinations of overriding defaults and defining runtime values were not representable before.
In terms of future enhancements, there is a lot we cannot track with the current models - such as map/reduce options for collection operations (https://github.com/galaxyproject/galaxy/issues/4623#issuecomment-389544980). This should enable a lot of that. Obviously there are a lot of attributes defined here that are not yet utilized, but I'm using most (all?) of them downstream in the CWL branch. I'd rather populate this table fully realized and fill in the implementation around it as work continues to stream in from the CWL branch - to keep things simple and avoid extra database migrations. But I understand if this feels like speculative complexity we want to avoid despite the implementation being readily available for inspection downstream.
Refactor get_data_inputs into get_all_inputs and use the resulting dictionaries to reason about if collection mapping should occur during invocation of tools. Using these dictionaries instead of explicit tool input objects should allow reuse within other module types since they may produce the same interface.
Fix warnings during py34-unit tests like:
```
/home/travis/build/galaxyproject/galaxy/test/unit/workflows/test_extract_summary.py:46: DeprecationWarning: Please use assertEqual instead.
self.assertEquals(job_dict[hda.job], [('out1', derived_hda_2)])
```
See e.g. https://travis-ci.org/galaxyproject/galaxy/jobs/438829277