Commit Graph
48 Commits
Author SHA1 Message Date
Nate Coraor 9aed5a36d8 Fix for dataset cleanup scripts and object store. 2012-01-31 14:07:14 -05:00
Nate Coraor 3a1526ac8a Decrease user disk usage when files are removed from disk by the cleanup_datasets.py. 2011-08-22 16:19:49 -04:00
Greg Von Kuster 581c3d9c67 Another patch from Assaf Gordon for the cleanup_datasets.py script.
This patch adds messages for the "purge histories" step:
1. A message is printed when processing a new history (so that the user can tell to which history the following dataset/hda messages belong).
2. in "info_only" mode, a message is printed about each deleted history
3. in "info_only" mode, a message is printed for each HDA that is checked
4. The message for HDA is changed to include the associated dataset id.
2011-07-14 15:58:06 -04:00
Greg Von Kuster 8eb0033904 Add patch from Assaf Gordon to the cleanup_datasets.py script. The changes are:
1. If a dataset is skipped (because it's shared/cloned and was already process), no message is printed at all.
2. If a dataset can not be deleted because it is shared, and one instance is not marked as "deleted", a proper message is printed.
3. If a dataset has metadata files, the message is changed depending on "info_only" and "remove_from_disk" flags.
4. The final summary message is slightly changed.
2011-07-13 09:39:02 -04:00
Greg Von Kuster 40d02d1e7b Add patch from Assaf Gordon to fix the cleanup_dataset.py script, correcting the behavior of setting the info_only flag to True when running delete_datasets.sh and purge_datasets.sh. 2011-07-06 09:52:28 -04:00
Ross Lazarus d53c3e984f Backed out changeset 48bbe32beefe which introduced a whole bunch of unintended reversions from a broken hg repository
This is a backout of commit 5765
2011-07-06 09:44:56 +10:00
Ross Lazarus 86b55bb0c0 branch merge 2011-07-05 12:40:43 +10:00
Greg Von Kuster b0b5df4e93 Add a purged column to the LibraryDataset table - a LibraryDataset is marked purged when all associated LibraryDatasetDatasetAssociations are marked deleted. Fix the cleanup_datasets.py script to more correclty handle the lifecycle of LibraryDatasets. 2011-02-09 15:42:09 -05:00
Kanwei Li 7736a34123 Add default file_path to cleanup_datasets.py script. Closes #376 2010-10-07 14:15:19 -04:00
Greg Von Kuster 91d97585c1 Fix for cleanup_datasets.py - add app param to the _purge_dataset method signature. 2009-11-14 20:32:41 -05:00
Greg Von Kuster 8f6e8213a0 Fix most of the db flushes to be compatibel with sqlalchemy 05. Add the _monkeypatch_query_method() back into assignmapper due to a single object.get() method in the MetadataCollection class since Metadata has no current hook into mapping.context ( the sqlalchemy session ). There a 4 flushes in metadata,py and 20 flushes in model.__init__.py that still use the _monkeypatch_session_method in assignmapper due to the same issue, but all other flushes are fixed. 2009-11-11 15:59:41 -05:00
Greg Von Kuster 1c4070ea92 Eliminate printing PYTHONPATH to stderr in the cleanup_datasets.py script. 2009-11-04 13:09:37 -05:00
Greg Von Kuster da149187b4 Bug fix for cleanup_datasets.py - use app.sa_session instead of self.sa_session. 2009-11-02 09:54:18 -05:00
Greg Von Kuster 29b55b8707 Make sqlalchemy queries use sa_session.query( object ) rather than object.query(). This eliminates the need for the _monkeypatch_query_method() in assignmapper.py. Also eliminate the need for the _monkeypatch_session_method() for everything except object.flush() - I just ran out of time and will handle flush() asap. I also eliminated the .all() call on several queries where it was not necessary - should improve performance. I fixed several bugs I found as well. 2009-10-22 23:02:28 -04:00
Greg Von Kuster 18c95bb622 Improve logging for cleanup_datasets.py script. 2009-09-24 21:17:32 -04:00
Greg Von Kuster 947491b270 More improvement in cleanup_datasets script - will take significantly less disk, and may be faster. 2009-09-24 16:52:15 -04:00
Greg Von Kuster a40bf886f1 Performance improvements in delete_datasets method in cleanup_datasets script ( hopefully enough ). 2009-09-24 12:21:41 -04:00
Nate Coraor 2abf463a95 More fixes for cleanup_datasets.py 2009-09-15 12:57:17 -04:00
Nate Coraor f11dcd9781 Fix path mangling in cleanup_datasets.py 2009-09-14 12:23:16 -04:00
Greg Von Kuster 5665f1dbaf Fix for cleanup_datasets.py script. 2009-09-11 11:26:26 -04:00
Daniel Blankenberg 52d1bc065a Enable the cleanup_datasets script to purge Dataset Instances (HDAs, LDDAs) and mark base datasets as deleted without requiring the container (history, library/folder) to be deleted. 2009-08-10 11:53:28 -04:00
Daniel Blankenberg f4a7dbc436 Add a new flag -f/--force_retry to cleanup_datasets.py. This flag will cause the script to attempt to perform the requestion action on objects regardless if it has been performed before.
This is useful, i.e. if purge_datasets was called with out using --remove_from_disk, but it is later decided to remove these files: the purge_datasets script should be called with both -r and -f.
2009-07-30 12:56:16 -04:00
Greg Von Kuster 0090622f59 Add ability to share histories with multiple users, along with bug fixes and more functional test coverage for history features. 2009-06-05 11:15:25 -04:00
Nate Coraor 9113ceb29f Fix the cleanup_datasets script 2009-05-28 15:18:49 -04:00
Daniel Blankenberg da52bd848a Add ability for admins to delete and undelete library items.
Modify cleanup_datasets.py script and include a database migration script.
2009-04-23 12:42:05 -04:00
Greg Von Kuster cfde4c4015 Normalize the model supporting the new library features. This rev assumes all current Library tables will be dropped so they are re-created at server start-up ( all library data will be lost ). Normalized objects are as follows:
1.  ActionDatasetRoleAssociation -> DatasetPermissions ( renamed )
2.  LibraryFolderDatasetAssociation -> LibraryDatasetDatasetAssociation ( renamed )
3.  ActionLibraryItemRoleAssociation -> LibraryPermissions, LibraryFolderPermissions, LibraryDatasetPermissions, LibraryDatasetDatasetPermissions, LibraryItemInfoPermissions, LibraryItemInfoTemplatePermissions
4.  LibraryItemInfoAssociation -> LibraryInfoAssociation, LibraryFolderInfoAssociation, LibraryDatasetInfoAssociation, LibraryDatasetDatasetInfoAssociation
5. LibraryItemInfoTemplateAssociation -> LibraryInfoTemplateAssociation, LibraryFolderInfoTemplateAssociation, LibraryDatasetInfoTemplateAssociation, LibraryDatasetDatasetInfoTemplateAssociation
2009-01-28 16:29:30 -05:00
Greg Von Kuster 2dd60568fc merging from central 2008-12-01 14:07:33 -05:00
Greg Von Kuster 4015f31eff Refix cleanup_dataset.py script. 2008-12-01 14:07:36 -05:00
Greg Von Kuster 5d40b2b545 merging from central 2008-11-30 08:05:04 -05:00
Greg Von Kuster 3748e41924 Fixes, code cleanup in cleanup_datasets.py. 2008-11-30 08:03:21 -05:00
Greg Von Kuster 34cda87570 Add ability to manage deleted libraries, requires db schema change, commands are:
ALTER TABLE library ADD COLUMN purged BOOLEAN DEFAULT FALSE;
CREATE INDEX ix_library_purged ON library USING btree (purged);

ALTER TABLE library_folder ADD COLUMN purged BOOLEAN DEFAULT FALSE;
CREATE INDEX ix_library_folder_purged ON library_folder USING btree (purged);

- Deleting a library will mark all folders and contents ( LibraryFolderDatasetAssociations ) deleted in the db (weakness is that state is not saved for undeleting )
- Undeleting a library will mark all folders and contents as undeleted in the db
- Purging a library will mark the library, all folders, and LibraryFolderDatasetAssociations as purged, datasets will only be marked as deleted - marking as purged and removing the file from disk will be handled by the cleanup_datasets script.
2008-11-12 14:35:07 -05:00
Greg Von Kuster e0df0fa44e merging from central 2008-11-10 15:30:20 -05:00
Greg Von Kuster ea7074e280 Fix for purging dataset - add check for pre-history_dataset_association approach to sharing. 2008-11-10 15:30:07 -05:00
Daniel Blankenberg ec188a33cb Resolve a bunch of merge conflicts. 2008-10-31 11:28:09 -04:00
Greg Von Kuster 5be8fefaf5 Migrate central repo to alchemy 4. 2008-10-30 16:17:46 -04:00
Greg Von Kuster 3b78d585c0 Purge metadata files associated with a purged dataset via the library association. 2008-10-28 14:44:09 -04:00
Greg Von Kuster 1f3494adc3 merging from central 2008-10-28 14:31:36 -04:00
Greg Von Kuster 9a7bc1370d Purge metadata files associated with a dataset when the dataset is purged. Also remembered log.exception logs the exception, so corrected a few things in jobs.__init__. 2008-10-28 14:31:02 -04:00
Greg Von Kuster 8e17384a79 Upgrade to SQLAlchemy 0.4.7 and correct conflicts in ~/model/__init__.py. 2008-09-18 15:18:12 -04:00
Greg Von Kuster 3de877519e Migrate cleanup_datasets script to work with new db schema. 2008-07-10 19:54:46 +00:00
Daniel Blankenberg 28b89d686f Change database schema to separate Dataset and HistoryDatasetAssociation. This is a significant change, be sure to follow the migration notes exactly.
Make sure to backup your database before updating.


Also fix a long standing bug in the async controller dealing with updating output datasets upon job completion.


Credit for database migration directions go to Greg.

*** This is a significant change, be sure to follow the migration steps below exactly and in order: ***

---------------

0. Stop Galaxy server

---------------

1. Backup database and galaxy_root/database directory

---------------

2. The current validation_error table is not being used, it is left over from Ian's work that was not made functional.  However, if we want to keep it around, we should drop the old version of the table so it will get re-created correctly when the Galaxy server is restarted:

any database - sql command(s):
------------------------------
DROP TABLE validation_error;

---------------

3. Stop server, perform svn update to get latest code, start server to create new history_dataset_association and implicitly_converted_dataset_association tables

---------------

4. Add new columns to dataset table:

postgres sql command(s):
------------------------
ALTER TABLE dataset ADD COLUMN purgable boolean DEFAULT 't';
ALTER TABLE dataset ADD COLUMN external_filename text;
ALTER TABLE dataset ADD COLUMN _extra_files_path text;

mysql sql command(s):
---------------------
ALTER TABLE dataset ADD COLUMN purgable boolean DEFAULT TRUE;
ALTER TABLE dataset ADD COLUMN external_filename text;
ALTER TABLE dataset ADD COLUMN _extra_files_path text;

sqlite sql command(s):
----------------------
ALTER TABLE dataset ADD COLUMN purgable boolean DEFAULT 0;
ALTER TABLE dataset ADD COLUMN external_filename text;
ALTER TABLE dataset ADD COLUMN _extra_files_path text;

---------------

5. Populate new columns:

postgres sql command(s):
------------------------
UPDATE
dataset
SET
external_filename = dataset_filename.filename,
_extra_files_path = dataset_filename.extra_files_path
FROM
dataset_filename
WHERE
dataset.filename_id = dataset_filename.id;

mysql or sqlite sql command(s):
-------------------------------
UPDATE
dataset,
dataset_filename
SET
dataset.external_filename = dataset_filename.filename,
dataset._extra_files_path = dataset_filename.extra_files_path
WHERE
dataset.filename_id = dataset_filename.id;

---------------

6. Populate dataset table with info from dataset_child_association table:

postgres sql command(s):
------------------------
UPDATE
dataset
SET
parent_id = dataset_child_association.parent_dataset_id
FROM
dataset_child_association
WHERE
dataset.id = dataset_child_association.child_dataset_id;

mysql or sqlite sql command(s):
-------------------------------
UPDATE
dataset,
dataset_child_association
SET
dataset.parent_id = dataset_child_association.parent_dataset_id
WHERE
dataset.id = dataset_child_association.child_dataset_id;

---------------

7. Drop dataset_child_association table:

any database - sql command(s):
------------------------------
DROP TABLE dataset_child_association;

---------------

8. Copy parts of the current dataset table to the new history_dataset_association table:

postgres sql command(s):
------------------------
INSERT INTO history_dataset_association
SELECT
id,
history_id,
id AS dataset_id,
create_time,
update_time,
hid,
name,
info,
blurb,
peek,
extension,
metadata,
parent_id,
designation,
deleted,
visible
FROM dataset
ORDER BY id;

mysql or sqlite sql command(s):
-------------------------------
INSERT INTO
history_dataset_association
(
history_id,
dataset_id,
create_time,
update_time,
hid,
name,
info,
blurb,
peek,
extension,
metadata,
parent_id,
designation,
deleted,
visible
)
SELECT
history_id,
id,
create_time,
update_time,
hid,
name,
info,
blurb,
peek,
extension,
metadata,
parent_id,
designation,
deleted,
visible
FROM
dataset
ORDER BY
id;

---------------

9. Update the nextval value of the primary key sequence in the history_dataset_association table

NOTE: THIS IS NOT NECESSARY FOR MYSQL OR SQLITE!

postgres sql command(s):
------------------------
SELECT
setval('history_dataset_association_id_seq', max(id))
FROM
history_dataset_association;

---------------

10. Update dataset table to mark dataset_associated_file dataset as deleted:

postgres sql command(s):
------------------------
UPDATE dataset
SET
deleted = 't'
WHERE
id IN
(SELECT dataset_id FROM dataset_associated_file ORDER BY dataset_id DESC);

mysql or sqlite sql command(s):
-------------------------------
UPDATE dataset
SET
deleted = TRUE
WHERE
id IN
(SELECT dataset_id FROM dataset_associated_file ORDER BY dataset_id DESC);

---------------

11. Drop the dataset_associated_file table:

any database - sql command(s):
------------------------------
DROP TABLE dataset_associated_file;

---------------

12. Alter foreign keys on job_to_input_dataset table:

postgres or sqlite sql command(s):
------------------------
ALTER TABLE
job_to_input_dataset
DROP CONSTRAINT
job_to_input_dataset_dataset_id_fkey;

ALTER TABLE
job_to_input_dataset
ADD CONSTRAINT job_to_input_dataset_dataset_id_fkey
FOREIGN KEY
(dataset_id)
REFERENCES
history_dataset_association(id);


mysql sql command(s):
-------------------------------
ALTER TABLE
job_to_input_dataset
DROP FOREIGN KEY
ix_job_to_input_dataset_dataset_id;

ALTER TABLE
job_to_input_dataset
ADD FOREIGN KEY
ix_job_to_input_dataset_dataset_id (dataset_id)
REFERENCES
history_dataset_association(id);

---------------

13. Alter foreign keys on job_to_output_dataset table:

postgres or sqlite sql command(s):
------------------------
ALTER TABLE
job_to_output_dataset
DROP CONSTRAINT
job_to_output_dataset_dataset_id_fkey;

ALTER TABLE
job_to_output_dataset
ADD CONSTRAINT
job_to_output_dataset_dataset_id_fkey
FOREIGN KEY (dataset_id)
REFERENCES history_dataset_association(id);

mysql sql command(s):
-------------------------------
ALTER TABLE
job_to_output_dataset
DROP FOREIGN KEY
ix_job_to_output_dataset_dataset_id;

ALTER TABLE
job_to_output_dataset
ADD FOREIGN KEY
ix_job_to_output_dataset_dataset_id
FOREIGN KEY (dataset_id)
REFERENCES history_dataset_association(id);

---------------

14. Eliminate columns from the dataset table previously copied to history_dataset_association:

any database - sql command(s):
------------------------------
ALTER TABLE dataset DROP COLUMN hid;
ALTER TABLE dataset DROP COLUMN history_id;
ALTER TABLE dataset DROP COLUMN name;
ALTER TABLE dataset DROP COLUMN info;
ALTER TABLE dataset DROP COLUMN blurb;
ALTER TABLE dataset DROP COLUMN peek;
ALTER TABLE dataset DROP COLUMN extension;
ALTER TABLE dataset DROP COLUMN dbkey;
ALTER TABLE dataset DROP COLUMN metadata;
ALTER TABLE dataset DROP COLUMN parent_id;
ALTER TABLE dataset DROP COLUMN designation;
ALTER TABLE dataset DROP COLUMN visible;
ALTER TABLE dataset DROP COLUMN filename_id;
2008-07-09 17:41:08 +00:00
Greg Von Kuster aa66fa9a61 Missed a place in cleanup_datasets.py where I wanted to turn off logging. 2008-06-27 13:16:36 +00:00
Greg Von Kuster 55cbdf9905 Tweak to cleanup_datasets script, only log error messages. 2008-06-20 13:45:38 +00:00
Greg Von Kuster 94917f5fe0 Added some important info to cleanup_datasets.py script. 2008-05-15 20:40:19 +00:00
Daniel Blankenberg a7ea9357f2 Add implicit datatype conversions.
If a dataset is not the proper format required by a tool, but there is a converter available, the dataset will appear as a valid option in DataToolParameters.

The converted dataset will be created (and subsequently reused) when the job is executed. Conversion utilizes the job runner, and will run on the cluster.

Setting metadata will invalidate the converted dataset, and a new one will be automatically generated as needed.
2008-05-15 19:09:17 +00:00
Greg Von Kuster 3f01a6dd2a Necessary cron changes for cleanup dataset scripts for main. 2008-04-24 20:31:27 +00:00
Nate Coraor 1d5f59a3c3 Updated dataset cleanup scripts to use eggs lib. 2008-04-18 13:53:15 +00:00
Greg Von Kuster a9db5f0e43 Moved all dataset related shell and python scripts to new ~/<galaxy root>/scripts/cleanup_datasets directory. 2008-04-03 18:33:09 +00:00