There used to be a long-standing bug in setting up the MetadataValidator
if default values were used. We fixed this in
https://github.com/galaxyproject/galaxy/pull/13139/commits/e194ef97e49c4947711ad1d20873ddc0aa686a66,
but that means we're now checking for all non-optional values before
running a tool. We have a ton of non-optional MetadatElement items
in datatypes that should maybe be optional (an indication might be if
`default` and `no_value` are specified and set to the same value ... but
I'm not sure that's a 100% thing). So I think that reviewing this
requires domain knowledge of the datatypes and what elements are really required,
and I'm not sure we can do this in a timely fashion, and not break
something that used to work.
So my suggestion is that we add `check_required_metadata=True`
on datatypes for which we have checked that non-optional metadata
elements are really non-optional. For those metadata elements
for which this is not the case we skip the validation as we would
do prior to
https://github.com/galaxyproject/galaxy/pull/13139/commits/e194ef97e49c4947711ad1d20873ddc0aa686a66.
As an example I have marked RDS and RData with check_required_metadata
and added a test for check_required_metadata.
Fixes https://github.com/galaxyproject/galaxy/issues/13090:
```
Sun, Dec 19 2021 9:01:49 pm | Traceback (most recent call last):
Sun, Dec 19 2021 9:01:49 pm | File "/galaxy/server/lib/galaxy/model/store/discover.py", line 203, in set_datasets_metadata
Sun, Dec 19 2021 9:01:49 pm | primary_data.set_meta()
Sun, Dec 19 2021 9:01:49 pm | File "/galaxy/server/lib/galaxy/model/__init__.py", line 3736, in set_meta
Sun, Dec 19 2021 9:01:49 pm | return self.datatype.set_meta(self, **kwd)
Sun, Dec 19 2021 9:01:49 pm | File "/galaxy/server/lib/galaxy/datatypes/interval.py", line 412, in set_meta
Sun, Dec 19 2021 9:01:49 pm | Tabular.set_meta(self, dataset, overwrite=overwrite, skip=i)
Sun, Dec 19 2021 9:01:49 pm | File "/galaxy/server/lib/galaxy/datatypes/tabular.py", line 404, in set_meta
Sun, Dec 19 2021 9:01:49 pm | if dataset_fh.tell() != dataset.get_size():
Sun, Dec 19 2021 9:01:49 pm | OSError: telling position disabled by next() call
```
Include a unit test that hits this error.
Previously sep2tabs would imply to_posix_lines even if that was explicitly set to false. I think this new behavior is what I intended and is more correct.
Here we are supposed to be testing mapping; however, by using Galaxy's
session and engine, we also end up testing session and engine setup
(i.e., "testing by coincidence"). Instead, we will use a plain engine
and session, thus limiting these tests to their expected scope.
Applied to model and install_model mapping tests.
It is possible (even likely, with a different/better/simpler session
setup) that the objects populating the collection attribute representing
a relationship on a stored object will NOT be the same Python objects
that were initially associated with the object via that relationship.
Thus, Python object identity comparison will fail (when we change the
session setup). Besides, we should be using the database primary key as
the comparison key, because that is the only correct measure of
database-mapped object equality.
Rename metadata to metadata_. This will break things.
(we are only renaming the model attribute, not the table column)
Next step: rename all instances of [ToolShedRepository].metadata in code base.
- Create a new project galaxy-files that does not depend on anything but galaxy-util.
- Add a dependency in galaxy-data on galaxy-files, use typing in galaxy.data.sniff to make this clear.
- Cleanup the galaxy-files unit tests to not depend on galaxy.datatypes.sniff code - this wasn't ideal to start with.
- Add a small integration-ish unit test test_sniff_file_sources.py to test that piece of sniff.py that depends on galaxy-files.