· 9 years ago · Apr 21, 2017, 04:32 PM
1* c945c7b (HEAD -> master) after constandpp packageification: refactoring: scriptsTools.py->tools.py refactoring all import statements: file->constandpp.file
2* 73be839 filled setup.py correctly
3* 25ecad7 (origin/master, origin/HEAD) undo silly file creations due to subtree magic and then merging with origin
4|\
5| * f20f42f (refs/original/refs/remotes/origin/master, refs/original/refs/heads/master) moved unnest() from dataIO.py to new file tools.py
6| * 1ffbc30 MAIN SCRIPTS CLEANUP: removed devStuff()
7| * 0ad971a MAIN SCRIPTS CLEANUP: moved, docu and refactoring of dataSuitabilityMA->testDataSuitability to scripts dir
8| * 08a630d MAIN SCRIPTS CLEANUP: moved, docu of compareDEAresults to scripts dir
9| * 8cbbacd MAIN SCRIPTS CLEANUP: moved, docu of intraInterMAPlots to scripts dir
10| * 30071a9 MAIN SCRIPTS CLEANUP: removed compareAbundancesIntSN()
11| * 72f9ef7 MAIN SCRIPTS CLEANUP: moved, docu and refactoring abundancesPCAHCD()->PDAbundancesPCAHCD to scripts dir
12| * 88637e6 MAIN SCRIPTS CLEANUP: moved, docu compareICmethods() to scripts dir
13| * cace51e MAIN SCRIPTS CLEANUP: moved, docu and refactoring: compareIntensitySN->testIntensitySNDifference to scripts dir
14| * 04ae43e moved fontsize and fontweight global variables from main.py to __init__.py
15| * 515ff4a MAIN SCRIPTS CLEANUP: moved MA(), scatterplot(), MAPlot(), boxPlot(), RDHPlot() to tools.py in the scripts dir
16| * 2bcc44e MAIN SCRIPTS CLEANUP: removed testDataComplementarity()
17| * bd1950d MAIN SCRIPTS CLEANUP: removed MS2IntensityDoesntMatter()
18| * 93cb8d6 MAIN SCRIPTS CLEANUP: moved and refactoring: isotopicCorrectionsTest()->testPDIsotopicCorrectionsEffect.py
19| * 0c7664c MAIN SCRIPTS CLEANUP: moved, docu and refactoring: isotopicCorrectionsTest()->testPDIsotopicCorrectionsEffect.py
20| * d5782e7 MAIN SCRIPTS CLEANUP: moved and refactoring: isotopicImpuritiesTest()->testIsotopicCorrectionNecessity.py
21| * c00d4a2 MAIN SCRIPTS CLEANUP: removed performaceTest()
22| * 4649b6b moved profiler.py from constandpp to scripts repo
23| * f5bf642 applyFoldChange() docu clarification
24| * 75530b1 smaller scatterplot fig size
25| * d1acadc fixed bug in scatterplot()
26| * c3385b3 added int64 as viable data type for constand
27| * 600f7c2 added default parameters to constand function
28| * 3ecf5f1 fixed small bug in scatterplot. Probably the MAPlot function won't work that properly anymore.
29| * 63b2e96 added scatterplot script function to main.py that MAPlot now depends upon
30| * 63669f4 changed emailadress to UHasselt address
31| * f9a2402 added __init__.py and moved credentials there
32| * 359690e fixed bug due to renaming
33| * de4405c fixed unused imports
34| * 3ce0bf0 fixed webFlow unused import
35| * ca78846 docu send_mail in web.py
36| * c357e0f deleted obsolete constructMasterConfigContents in web.py
37| * 6f45e86 line endings bullshit
38| * 02835ce refactoring modified TMT_ICM->TMT_IDT files in job folder stuff
39| * 96be4d1 refactoring modified TMT_ICM->TMT_IDT files in job folder stuff
40| * db8c6cb docu startJob in web.py
41| * 390a6da stuff
42| * e2528d2 fixed docu mistakes in setJobCompleted and setJobFailed
43| * e12719f docu DB_setJobCompleted in web.py and rename jobDirName->jobID
44| * 998e8c9 docu DB_setJobReportRelPaths in web.py and rename jobDirName->jobID
45| * a4f6fc3 docu makeJobConfigFile in web.py
46| * 0afc984 docu updateWrappers in web.py
47| * b30ae44 added todo because there is a file manipulation function location inconsistency between dataIO.py and web.py
48| * 8515d56 added jobs symlink in static folder to git ignore
49| * 200ea7c docu updateConfigs
50| * 90ae5ba docu DB_getJobVar in web.py
51| * c4b3a6c docu DB_checkJobExists and DB_insertJob in web.py
52| * 90af650 docu updateSchema in web.py
53| * 346711e docu saveFileStorage in web.py
54| * 75c4e1c docu hackExperimentNamesIntoForm
55| * 4ba574d docu web.py file description
56| * 735e8b4 docu newJobDir in web.py
57| * 1abfc24 webFlow.py docu file description
58| * 76f5739 docu hackImagePathToSymlinkInStaticDir
59| * 6cf0afa created injectColumnWidthHTML() in report.py to fix the relative column widths in the pandas-generated differentials table
60| * ee9fc4d used the pandas to_html() function to generate differential proteins table HTML instead of trying to loop over a dataframe in Jinja
61| * b447cae stuff
62| * 6dbf774 fixed obsolete argument bug in getPCA call to getMarkers
63| * fd8b0de added job failure message to jobInfo if job is done but failed
64| * 6e78fb6 put try-except around mailer.send(msg) in send_mail() in web.py and used a flash message to inform the user via the web. So now the application des not stop if a mail doesnt get sent
65| * 44ed327 removed obsolete code in views.py
66| * 598f03f fixed bug in jobInfo() in views.py if the ID was invalid. Updated the error message to contain link back to /jobinfo
67| * 8d68ed7 docu jobInfo in views.py
68| * 59c4c52 moved allJobsDir and jobDB parameters from __init__.py jo config.py and made uppercase
69| * 10abbe8 updated config.py
70| * 0ed567f added __pycache__ to gitignore
71| * ad51282 editing: jobInfo docu in views.py and parameters in config.py
72| * d768ce7 docu jobSettings in views.py
73| * 827e5cd docu up to jobSettingsForm in views.py
74| * 3631987 docu __init__.py about the views import statement at the end of the file
75| * 62e488a docu forms.py
76| * 3a07139 docu config.py
77| * 4272362 docu __init__.py and rename of parameter DB->jobDB
78| * d8acd59 docu collapse() in collapse.py and removed obsolete code
79| * 62c72f2 docu getRepresentativesDf in collapse.py
80| * 542ee42 docu getIntenseIndicesDict in collapse.py
81| * 6d1b8f7 docu getBestIndicesDict in collapse.py
82| * ed81fab docu combineDetections in collapse.py
83| * dceeb46 docu groupByIdenticalProperties
84| * 7fda591 refactoring: getData->importExperimentData
85| * c57e8ac refactoring: getWrapper->importWrapper
86| * a42c28d refactoring: getIsotopicCorrectionsMatrix->importIsotopicCorrectionsMatrix
87| * 22a558e added .gitignore file after git-extras package installation
88| * c0b6c88 docu dataIO
89| * 8926941 main.py fixed some layout and comment details
90| * c4dee43 commented out writeConfig() + docu (partial) dataio.py
91| * a9482a5 refactoring masterConfig->jobConfig and docu getinput.py
92| * 908f185 docu runweb.py
93| * 79bcc46 docu makeHTML in report.py
94| * ad81c31 docu HTMLtoPDF in report.py
95| * d0eda62 docu getHCDendrogram in report.py
96| * 0053d52 docu and removed obsolete code getPCAPlot in report.py
97| * 5fb575b docu getVolcanoPlot
98| * ce3945e docu and removed obsolete argument getMarkers() in report.py
99| * fc6db66 docu getColours
100| * acf8a54 convert indents to tabs
101| * 57c83e3 docu getColours
102| * 5ddc023 docu generateReport in reportFlow.py
103| * 190ec22 fixed main.py imports
104| * 996d63b (refs/original/refs/remotes/origin/webinterface, refs/original/refs/heads/webinterface) MERGE UPDATES FROM BROKEN TO ORIGINAL REPO started docu reportFlow.py
105| * 9fd790e MERGE UPDATES FROM BROKEN TO ORIGINAL REPO fixed imports in main.py
106| * fc30358 MERGE UPDATES FROM BROKEN TO ORIGINAL REPO docu processing.py
107| * 9e7f92b MERGE UPDATES FROM BROKEN TO ORIGINAL REPO docu processingFlow.py
108| * de27b72 MERGE UPDATES FROM BROKEN TO ORIGINAL REPO docu analysis.py
109| * e488fa2 MERGE UPDATES FROM BROKEN TO ORIGINAL REPO refactor applyDifferentialExpression->testDifferentialExpression docu testDifferentialExpression
110| * 56c90a0 put all paths in app config
111| * 60e8a9b Merge branch 'webinterface' of ssh1.ulyssis.org:thesis into webinterface
112| |\
113| | * feeb0ac moved updateConfigs to web.py; fixed some bugs
114| * | e7419ac set up mail settings
115| * | 90cf212 setup mail attachments
116| * | 4e3f057 added email to forms
117| * | 990f134 set autorefresh to 5s
118| * | af0053c fixed the fucking form bugs satan will have my soul for this dirty code
119| * | 880a303 fixed bug url_for image generation bug by just FUCKING NOT USING IT (nad not setting SERVER_NAME because it breaks everything)
120| * | 64f70a3 bug in url_for usage in report.html when SERVER_NAME is not set
121| * | 0acd46b bugfix: getFile() view should use request args, not url argument
122| * | 426c388 weird bug in analysisFlow for a missing parameter that should have surfaced 1 month ago
123| * | b648686 bugfixes in volcanoFullPath declaration (now always defined, sometimes as None)
124| * | 9161c30 implemented makeHTML() and HTMLtoPDF
125| * | d0a3caa added some metadata
126| * | fe65e7b exporting figures now returns their full path
127| * | f27f49d made report template report.html
128| * | 7025630 HTML and PDF reports now show via jobInfo()
129| * | 26a2199 fixed jobInfo refresh, now working on showing the HTML and PDF reports
130| * | 6819d38 FINALLY fixed bug where the program didn't run
131| * | 12500d0 fixed report generation flow skeleton
132| * | 3becaee startJob now returns Popen object
133| * | 57342ad bugfixes
134| * | 610d5e9 added delim_out to jobSettings form, fixed app context bug (?)
135| * | 476b809 fixed some busg (mostly in startJob)
136| * | b889bf5 fixed isDone in jobInfo
137| * | 04dc6e8 removed DB_close calls
138| * | 5b89d02 added empty isotopicCorrection_matrix entry in updateSchema
139| * | 151edaf fixed BASEPROCESSINGCONFIG file path in config
140| * | 9b70c71 modified startJob to set stdin,stdout,stderr=None
141| * | 833a1e6 implemented startJob
142| * | 0b87dcc completed webFlow transfer to jobSettings() in views.py; implemented DB_setJobReportRelPaths, DB_setJobCompleted, DB_setJobFailed
143| * | 959aed3 refactoring in webFlow.py: master...->job...
144| * | 28bb3d1 moved updateConfigs and updateWrappers to web.py
145| |/
146| * 10b6ed0 implemented DB_checkJobExist(), DB_close(), DB_insertJob(), getJobVar() in web.py
147| * 31006fb fixed updateSchema; implementing what should be done after updateSchema
148| * 7fe5ce2 FINALLY fixed experiment schema labels :D by adding hackExperimentNamesIntoForm
149| * ac95f1c still fixing experiment schema labels
150| * d931f0e filled jobSettingsForm with more options
151| * 1e57e7e implemented jobInfo with getHtmlReport and getPdfReport
152| * e073874 added jobstartedMail
153| * ea2da78 included html save procedure in dataIO
154| * 9b52a04 added sqlite3 database use
155| * 915fcea implementing jobSettings forms
156| * 4f7e50a fixed schema upload added newjob_content.html to be included in newjob.html and index.html
157| * 5ac2bc7 added config file
158| * a8a28eb created jobSettings(), which gets called after successful schema upload.
159| * cdb46e3 renamed schemaForm to newJobForm moved newJobDir() from webFlow.py to web.py implementing schema file check after upload
160| * c35acc8 made schemaForm()
161| * 3b6e5aa added newjob.html and included it in index.html also added forms.py
162| * 76c61e0 added base.html and documentation.html and now uses jinja block extentions
163| * c2d7e47 added css added completed.html
164| * 0056c3e implemented mailer (untested) implementing css
165| * 9740673 implemented getFile for rendering images (or files) from different locations
166| * 87f1029 created link to dummy report file
167| * 0e05ac6 disabled plot show because it disallows the program to end without closing the figures
168| * e1b84d8 moved web.py and webFlow.py to web folder and fixed imports
169| * 703e8aa created web subdir with views.py
170| * 975f1ad removed tests
171| * c1b0ba0 Merge branch 'master' of ssh1.ulyssis.org:thesis
172| |\
173| | * 800171d removed m files
174| * | 5ee46cd removed m files and other obsolete stuff like tests
175| |/
176| * 12727ab adapted runweb.py to be open for any host
177| * df10c80 added runweb.py to run a web interface
178| * e0d5e60 split off constructMasterConfigContents and TMT2ICM from webFlow.py into web.py
179| * acdf763 volcano plot limits
180| * 9e12ab2 volcano plot limits
181| * 6506568 stuff remaking plots
182| * 94f7a21 PCA and volcano plot label size is 20 and marker size is now 80 and 160 in legend
183| * f425eff compareICmethods() improved
184| * 67379b7 MA plot improvement font is still normal but can be adjusted through fontweight
185| * 877b0fc MA plot improvement
186| * 7b6df08 adapted siotopicCorrectionsTest
187| * f5db3d6 schema is now an OrderedDict, so that its items() order is decided by the order in which the keys were added, i.e. the order of the experiments in the schema file uploaded by the user. this is the best fix since sliced bread
188| * 2d62a34 PCA_components parameter is now hardcoded in the webflow to be 2
189| * 263cbde comments
190| * 61c3807 removedData are now saved in a separate removedData folder per processing job
191| * a42b7dc implemented numDifferentials parameter to control the amount of proteins in the DEA list of the report
192| * 4e6b4ae Merge branch 'master' of ssh0.ulyssis.org:thesis
193| |\
194| | * 4618c1c undoublePSMAlgo now checks whether there actually are enough PSMALgo's to undouble, even if undoublePSMAlgo_bool = true
195| * | 1e667f8 refactoring: minProteinDF_bool -> minExpression_bool; fullProteinDF_bool -> fullExpression_bool
196| * | b2489fa adapted analysis workflow to only calculate minProteinDF and fullProteinDF if specified by minProteinDF_bool and minProteinDF_bool adapted reportFlow as well
197| * | eb69995 removed todo
198| * | 1184e35 undoublePSMAlgo now checks whether there actually are enough PSMALgo's to undouble, even if undoublePSMAlgo_bool = true
199| |/
200| * 3f91278 changed ICM string in TMT style isotopic corrections table to IDT string
201| * df02b02 fontsize 30
202| * b361ea6 improced RDHplot to take a quantity argument specifying fidderence between what quantity
203| * 276046a MAX_abundances_noTAM hard code in compareDEAresults()
204| * 722ab35 updated fontsize lines (now refer to variable fontsize)
205| * e8ab0c1 implemented dataSuitabilityMA() to make MA plots and show that the data is unsuitable for constand
206| * 0246fc9 adjust font sizes everywhere
207| * f15ed24 added Delta V boxplots to intraInterMAPlots()
208| * 8807b46 refactoring the variable names in intraInterMAPlots() so that they remain distinguishable (different variable names for each case)
209| * 36560ec implemented boxPlot() for intraInterMAPlots
210| * 9b94f35 improvements for RHDPlot
211| * 720e15f bugfix in compareDEAresults(): now uses the regular NON-log fold change
212| * eb2c5a6 implemented RDHPlot() : relative difference histogram plot. measures |x-y|/max(x,y)
213| * 789461a implemented compareDEAresults() for making MA plots of similar DEA analyses results (p-value and fold change)
214| * 632feae updated compareICmethods()
215| * c623bc0 replaced MA plots mean by MAD (mean absolute deviation)
216| * 23161a3 Merge branch 'master' of ssh1.ulyssis.org:thesis
217| |\
218| | * 981b86e refactoring: split up intraInterMAPlots() into compareDIFFERENTconditions and compareIDENTICALconditions by means of 2 booleans
219| * | b0be719 refactoring: split up intraInterMAPlots() into compareDIFFERENTconditions and compareIDENTICALconditions by means of 2 booleans
220| |/
221| * 05147b5 added plot of last MA comparison to INTER and INTRA sections of intraInterMAPlots()
222| * 4045e42 implemented intraInterMAPlots
223| * 75a4207 refactoring: took the MA part out of MAPlot() and put into separate MA() function implementing intraInterMAPlots
224| * 6b15d4d refactoring: took the MA part out of MAPlot() and put into separate MA() function implemented
225| * 90c3c23 refactoring: getAllExperimentsIntensitiesPerCommonPeptide now returns a dataframe (with the intensity channels as headers) instead of a matrix
226| * 596fee1 bugfix in getProteinDF when df doesnt contain 'Protein Descriptions': now automatically added (dummy)
227| * 6f8a4f5 started implementation of intraInterMAPlots() for making all intra and all inter-experimental MA plots for the Max data set
228| * 2281995 bugfix in HCD: x and y labels were switched
229| * bcff26b cherry-pick merge from abundances branch: compareAbundancesIntSN() implemented, getRTIsolationInfo now returns None if "representative first scan" or "RT" columns are missing.
230| * cbdd0ef added hard code for COON_SN_nonormnoconstand
231| * 0376c4c added compareICmethods() for comparing the isotope corrections of proteome discoverer with constand++
232| * 90c8f5d TMT2ICM now returns the ICM with normalized rows.
233| * f5caf46 collapsePTM is now enabled again!
234| * 859ff27 added separate COON_nonormnoconstand hard code in webFlow
235| * f3d6de8 MAPlot title argument is optional
236| * 8d7975b remove log files from repo
237| * 5499acd getSortedDifferentialProteinsDF now sorts on adjusted p value instead of fold change, because p value already incorporates fold change.
238| * 7262fd1 applyDifferentialExpression now also remove proteins that have all nan values for a certain condition. Keep the removed ones in metadata
239| * 2cd839d applyDifferentialExpression now returns nan if p-value would be zero (one condition has all values missing)
240| * e897c7b docu
241| * 674efb8 protein descriptions are identical across multiple peptides per protein: select only the first in the list (i.e. do not make a list of identical descriptions)
242| * 35fbdee tons of stuff to make the MB_Bon dummy experiment working again
243| * 1925905 Merge branch 'master' of pow.ulyssis.org:thesis
244| |\
245| | * 9915485 adding job folder in scripts dir
246| * | 69bafbf generateReport() now also writes min and full sorted list of differentials to disk
247| |/
248| * aa33006 disabled a test boolean
249| * 27bb93f bugfix icm extention should not be in exportData call in transformICM()
250| * 0939595 pass channelNamesPerCondition to transformICM, not channelAliasesPerCondition
251| * 836be03 TMT2ICM docu: column order is changed
252| * f8e6f8e bugfix ICMFile indentation in updateSchema()
253| * 57fdb32 bugfix ICMFile job path
254| * 789ff96 bugfix if ICMFile would be None
255| * 7a5bf17 TMT2ICM now reorders the columns and rows according to the nested list of channelAliasesPerCondition
256| * 2066b02 implemented transformICM which takes a TMT ICM uploaded file and turns it into a real ICM and reorders the columns conform to the channelAliasesPerCondition by employing TMT2ICM
257| * 65a220f added COON_noISO exptype in webFlow()
258| * 90789c7 refactoring: moved TMT2ICM() and constructMasterConfigContents() from dataIO.py to webFlow.py
259| * 26f80db removed empty entries from incompleteSchemaDict in parseSchemaFile also: this function should NOT be moved to web
260| * e8f01af TMT2ICM now returns floats in the ICM instead of percentages
261| * f5b224a bugfix in TMT2ICM: some observedChannels to not exist when Nplex is a subset of e.g. 10plex
262| * df00416 bugfix in getTMTIsotopicDistributions: index is now of type str
263| * 28c4ac3 bugfix in getTMTIsotopicDistributions: modification must be inplace
264| * cf98e38 docu change of TMT isotope table format
265| * 81334a3 implemented getTMTIsotopicDistributions to get a TMT table from file in the right dataframe format as require by TMT2ICM.
266| * 9f4e88a implemented TMT2ICM for converting TMT isotope impurity distribution tables to isotopic correction matrices
267| * 37b4cfe added doConstand option to abdundancesPCAHCD
268| * 09fb5c0 implemented abdundancesPCAHCD function to analyze abundances of COON data instead of intensities or SNRs
269| * 775bab6 bugfix: MAPlot ignore nan and inf values when calculating var and mean
270| * 6015710 undo previous bogas bugfix; it was a stupid bracket
271| * d13335e small bugfix in np import dtype
272| * dc88e8a MA plot title (if provided) also gets mean M and var M
273| * db16285 added test code for making peptide-level MA plots
274| * 2a7513b MAPlot() now takes a title as optinoal argument
275| * 3464a42 allExperimentsIntensitiesPerCommonPeptide is now added to reportFlow return and is included in analysisProcessingResults(Dump) so that it can be aesily passed to reportFlow
276| * f03b367 added exportData for allExperimentsIntensitiesPerCommonPeptide so that we can do experiments or statistics on the peptide level
277| * bd7f5ec fixed newlines at end of files
278| * 2cc36b1 fixed volcano plot code line location in reportFlow (in/outside if condition)
279| * dfa425d bugfix: baseJobConfig was removedfrom scripts directory, it is back now
280| * e77e62c stuff
281| * ee27ca4 implemented functionality in compareIntensitySN to NOT do processing, but just constand only
282| * cc7ec4b bugfix import constand
283| * b0aa767 changed compareIntensitySN pickle file location
284| * 2cba12c compareIntensitySN can now take two dataframes as argument
285| * dd6c6ce refactoring: created processingFlow.py, analysisFlow.py, reportFlow.py and moved the functions processDf(), analyzeProcessingResults(), generateReport() from main.py to there. fixed imports
286| * 46d8c33 refactoring: main(): masterConfigFilePath->jobConfigFilePath
287| * 2ad4182 fixed devStuff compareIntensitiesSN() put test job files in job directory
288| * 43c210e wrote some code (and immediately commented out) to check if allChannelAliases order remained consistent between getAllExperimentsIntensitiesPerCommonPeptide and getPCAPlot
289| * ea52dcc TMT2ICM removed comments
290| * 06c8d16 getPCAPlot: added experiment label legend
291| * cbfe7a7 hotfix: suddenly i get errors for DEA and volcano plot stuff if the number of conditions is not 2 ... should have happened before.
292| * a0a07a2 comments
293| * 172d655 getMarkers now returns a dict with markers per channel. Each experiment still has its own unique marker
294| * d54eb9d getMarkers and getColours now get allChannelAliases as an argument instead of recalculating each time
295| * d32f777 moved informative console prints in main into the doX booleans
296| * 3204d77 figures should now be saved in larger size
297| * a02029e figures should now be saved maximized
298| * 434ac5b bugfix: visualizationsDict artifact
299| * f60c20e makeHTML() has been adapted to "no more visualizationDict"
300| * e75b694 generateReport doesn't use a visualizationDict anymore. Just uses separate variable for each plot
301| * 9d68e80 exportData for type dataViz replaced by type fig. Doesn't accept dicts anymore
302| * 6f40396 hardcoded COON_SN_norm data
303| * 757bd83 added informative console output strings in main
304| * 13b094c small bugfix when constand is not performed
305| * 1b9ce16 hardcoded COON_norm experiment code
306| * f73e3f7 implemented hardcoded switch in processDf to NOT use constand!
307| * 4ba3051 added COON_SN to webFlow
308| * fd2359d bugfix
309| * 7fc9403 adapted getPCAPlot to use the new channelColorsDict
310| * fa26ad9 getColours now returns a dict with 1 colour(=value) per channel(=key).
311| * ab45c5f fixed HCD colors
312| * 6312889 increased PCA marker size
313| * 1017fa0 apaprently some os.path's were still just path's????
314| * fa15d76 implemented exportData for type viz: figures are saved both as png and as pickle so you can reopen and modify
315| * 68baf0a implementing colors for getHCDendrogram changed distinguishableColours() cmap to jet
316| * ead7f9d getColours docu more explicit
317| * 5e9ae0c getHCDendrogram(): quality of life improvements + adding colours per condition
318| * 5d0e7b7 bugfix: removed Exception in getInput if the output path already exists, because I check this again in the main() method anyway.
319| * 907a044 removed obsolete commented code
320| * 2ab4bc2 moved getProcessingInput inside the same loop over the experiments as the doProcessing boolean. (the second loop was obsolete)
321| * 0b2f36d refactoring: masterParams -> jobParams
322| * 854ed1a refactoring: specificParams -> processingParams
323| * a3d59b8 implemented method to continue previous analysis: webFlow now takes argument previousjobdirName
324| * c0ec5f2 bugfix in analyzeProcessingResults if noCorrectionIndices does not exist because isotopicCorrection_bool == False
325| * c27a39d importDataFrame can now handle dtypes so that getWrapper -- CORRECTION: parseSchema() -- can specify "str" so as not to interpret 126 as a float 126.0, then to be cast to the wrong type of str value
326| * c88ef9b bugfix in updateConfigs and getProcessingInput: ICM can be None
327| * f8c0cb9 job name should be entered together with schema job path now also contains the job name
328| * 8985a48 webFlow now takes an exptype argument specifying which hardcoded experiments to use webFlow can now handle up to 4 hardcoded experiments added COON hard-code
329| * 91748d2 webFlow can now handle up to 4 hardcoded experiments added COON hard-code
330| * 607e679 made getProteinPeptidesDicts (and thus the whole workflow) independent of "# Protein Groups" column
331| * 25be513 implementing TMT2ICM (not finished yet)
332| * 89d10a4 refactoring artifact
333| * 2745786 implementing TMT2ICM converter (for web eventually) removed getIsotopicCorrectionsMatrix default path value
334| * ec26ef1 fixed resultsDumpFilename paths
335| * 44fcc61 refactoring: dataproc.py -> processing.py
336| * 14136ff refactoring: intensityColumnsPerCondition -> channelNamesPerCondition
337| * bec308e removed intensityColumnsPerCondition from processingConfig. only needs intensityColumns
338| * 17ee746 refactoring: masterConfig -> jobConfig; config -> processingConfig; baseConfig->baseProcessingConfig
339| * 8984a51 refactoring: getMasterInput -> getJobInput; getInput -> getProcessingInput
340| * f935211 changed all os.path.relpaths by os.path.abspaths in main()
341| * 06f7998 bugfix: webFlow now corretly works on absolute paths but writes relative paths to config files
342| * f2f96d8 bugfix: reverted accidental change of the masterConfig.ini in the scripts dir
343| * 75b3ec1 restructuring: path parameters in config.ini and masterConfig.ini should now be relative to the job path (specified by the parent path of the config file path) this has also been restructured in webFlow Also the hardcoded input paths in webflow are now gathered at the start of the file
344| * b266999 restructuring: path parameters in config.ini should now be relative to the job path (specified by the parent path of the config file path) this has also been restructured in webFlow Also the hardcoded input paths in webflow are now gathered at the start of the file
345| * 1bffa74 todos
346| * 1395f91 bugfix: some remaining filename_out replaced by jobname
347| * 1da9183 bugfix: warnedYet was never set to true in isotopicCorrection
348| * ec7e634 bugfix: added logfile to generateReport call
349| * 4896a08 bugfix: something went wrong in the commit from 2 commits ago: mean of empty slice warning was still showing. also removed some warn() artifacts and replaced by logging.warning()s
350| * 7bf7497 bugfix: maxRelativeReporterVariance artifacts
351| * 3eff2f2 found and silenced mean of empty slice spam source
352| * 54ecd6a removed maxRelativeReporterVariance and its flagging code which isnt being used
353| * f4fcacb bugfix: masterparams does not require eName
354| * 6a35233 moved the output dir generation back inside the experiment loop (analysis+result ones to respective other locations) but now correctly changes name on each loop --> each experiment processing has its own results subfolder
355| * a15c40e moved the output dir generation outside the experiment loop
356| * 503a654 todo
357| * ae255a6 removed channelAliasesPerCondition from the specific params
358| * f808035 bugfix: forgot to dumps channelAliasesPerCondition
359| * 5b401f1 small bugfix: syntax
360| * c4d4637 bugfix: updateConfigs() wrongly wrote the original intensityColumnsPerCondition to the config.ini file instead of the aliases
361| * 73b857e bugfix: removed debugging artifact, replaced remaining warn()s by logging.warning()s
362| * 0e7cd77 wrapper dataframe now imported as type str
363| * a25efa9 refactorign: dataIO.py: getDataFrame() -> getData()
364| * 7df0acf bugfix path
365| * d9561fd Merge branch 'master' into webFlow (added logging)
366| |\
367| | * 2018286 introduced logging replaces warnings.warn
368| * | f56cf4e added channelAliasesPerCondition to getInput()
369| * | 35d0a87 added different subpaths for output (processing, analysis) and results refactoring: filename_out -> jobname
370| * | ad76f5b implemented parseDelimiter() to also allow for None delimiters
371| * | ee713e7 delim_in is allowed to be None
372| * | 182fe67 bugfix: updateConfigs now correctly discerns writing strings and non-strings updateMasterConfig too
373| * | 4661cbf added removedDataInOneFile_bool, header_in, collapse_maxRelativeReporterVariance to baseConfig
374| * | 143a300 updated dataIO.py: importDataFrame() so that it can let pandas automatically detect delimiters
375| * | cfef750 refactoring icm -> isotopicCorrection_matrix file_in -> data
376| * | 24b3db5 fixed some wrong uses of write(), dumps() combinations
377| * | 23dee8d bugfix: forgot to write some newlines
378| * | bb3cd36 removed collapse_method from baseConfig
379| * | 63e7d08 removed custom schema definition artifact from main()
380| * | 977a0ab refactoring shadowing variables
381| * | 5fe92ca bugfix: return masterConfigFile (is already full path)
382| * | d9e03f6 Merge branch 'master' into webFlow
383| |\ \
384| | |/
385| | * 297f3c6 updated importDataFrame: now drops lines that contain no values
386| * | 898829c if no wrapper is uploaded, creates empty wrapper file to append to afterwards.
387| * | 382a37f implemented updateWrappers
388| * | 8bfe50c removed path_in from masterConfig and getInput
389| * | c900877 clarified webFlow steps
390| * | abc2a3f webFlow.py: made all paths full isntead of just filenames.
391| * | ab53ad9 bugfix indentation in updateConfigs(); better detection of [DEFAULT] line
392| * | fc0ca54 separated webFlow() into a different file webFlow.py
393| * | 0589c88 added a newline first when appending to config files fout.write('\n') # so you dont accidentally append to the last line
394| * | 6d49f04 bugfix: config files should be appended with option a, not option w (write)
395| * | a7902cc newJobDir now returns an absolute path
396| * | d36fab5 bugfix in date that is written to masterConfig
397| * | 7b09ad0 bugfix: dumps() does not write like config --> modify the write() operations so they write just like configparser
398| * | 72e0786 bugfix in removal of files when uploadSchema fails; parseSchema argument path; updateConfigs;
399| * | d2278d0 bugfix in removal of files when uploadSchema fails and parseSchema argument path
400| * | fbac388 Merge branch 'webFlow' of ssh1.ulyssis.org:thesis into webFlow
401| |\ \
402| | * | 989a9ab implementing webFlow in main.py to simulate a web tool providing input added baseConfig.ini added ../jobs folder
403| * | | 32c360d implementing webFlow: implemented rest of step 3 and also step 4.
404| * | | 8637460 implementing webFlow: put files in /jobs folder and implemented step 2 and part of 3 renamed baseConfig.ini
405| * | | 7904802 implementing webFlow in main.py to simulate a web tool providing input added baseConfig.ini added ../jobs folder
406| | |/
407| |/|
408| * | 1cf1538 bugfix: constructMasterConfigContents should now properly print tabs
409| |/
410| * 981e5b9 added masterInput parameter path_in
411| * e273894 BUG NOT FIXED: updated constructMasterConfigContents which now STILL DOESNT handle delimiters correctly (visible format)
412| * 346d1cf updated constructMasterConfigContents which now handles delimiters correctly (visible format)
413| * a4c8688 import bugfix
414| * d4e4c15 implemented constructMasterConfigContents
415| * f07bc3e parseSchemaFile refactoring and docu
416| * f756d49 implementing parseSchema
417| * f1310c6 multi-experiment bugfix in getAllExperimentsIntensitiesPerCommonPeptide
418| * 7c1889a stuff (implemented and commented out bad attempt at trying to make labels fit properly on the plot)
419| * 83e279e small bugfix
420| * 0768cee hard bugfix in distinguishableMarkers. Fuck, be careful when assigning variables: use .copy(). I was destroying the matplotlib Markers archive.
421| * 3c3bc29 bugfixes in distinguishableMarkers: use the markers KEYS instead of VALUES
422| * b71f5a1 getPCAPlot is now multi-experimental and has the channelAliases as labels and each condition has its own colour
423| * b2a1a68 small bugfix
424| * a703dae bugfix in getColours maybe I should just use human readable for loops
425| * 2bf9515 small bugfix in getColours
426| * 7b6fffd refactoring: moved main.py:unnest to dataIO.py:unnest
427| * 4c47197 reimplemented getColours
428| * 4f9eda5 bugfix in getColours and getMarkers
429| * 6ecf21b implemented getColours
430| * 7be276f implemented getMarkers
431| * 59cc3cf implemented metadata['commonNanValues']
432| * 0c1cb2b getAllExperimentsIntensitiesPerCommonPeptide() now also returns metadata['uncommonPeptides']
433| * f915369 refactoring in getAllExperimentsIntensitiesPerCommonPeptide(): pcadf -> peptidesDf
434| * 1a21e7a refactoring: getAllExperimentsIntensitiesPerCommonPeptide
435| * 2647b4f fixed getPCA so that NaN values are set to zero 0.
436| * 52a94f1 refactoring fix
437| * a2198f3 implementing getAllExperimentIntensitiesPerPeptide
438| * a957207 implementing getPCAPlot for multiple experiments ... but first need to do analysis.py:getPCA for multiple experiments
439| * b364825 implemented distMarkers() for generating distinguishable markers
440| * b9dc12f added distColours to generate distinguishable colours
441| * 9628df0 getPCAPlot refactoring: now uses schema instead of intensityColumnsPerCondition
442| * 36debf0 stuff, bugfixes
443| * dc91a47 added wrapper to schema
444| * 2160001 removeMissing does not use "Quan Info" columns anymore to remove detections without quan values
445| * 1183a24 corrected removeObsoleteColumns docu.
446| * 2506b9d ease of use
447| * cfd8c00 implemented use of wrapper file wrapper.tsv moved input parameter modification getIsotopicCorrectionsMatrix and getWrapper away from the initial declaration of the variables (so that one first obtains them as a path string) commented out ICM checks and numpy import
448| * 66a96f3 refactoring: Identifying Node -> ~ Type implemented fix in fixFixableFormatMistakes if one wasnt using Identifying Node Type yet.
449| * ce32378 verified that flanking amino acids correction in fixFixableFormatMistakes works
450| * 084d08b small bugfix: allChannelAliases definition
451| * 08ee742 bugfix: cannot call .extend on [] expression. Just use "+" operator
452| * 8c8ca56 bugfix: conditionXIntensities should be Series
453| * f909e46 small bugfix and undid some previous changes
454| * 6b80c01 fixed getProteinDF() to properly handle multiple experiments. Now contains extra aggregation per experiment. now also interprets peptideIndices as MultiIndex
455| * f1c02be fixed getProteinDF() to properly handle multiple experiments. Now contains extra aggregation per experiment.
456| * b9a436f bugfix in proteinDF definition in getProteinDF()
457| * 1c29395 bugfix in analyzeProcessingResult: definition of noCorrectionIndices
458| * d7f4e15 global variables are giving me shit. Commenting them all out. fixing the dependencies
459| * 623014b global variables are giving me shit. Commenting them all out.
460| * c114bed fixed bug in getIntensities that was caused by a debugging artifact
461| * e3b3ea0 fixed bug in config2.ini intensityColumns
462| * 5eef2e1 made 2 separate config files for separate experiments (on same data file)
463| * f85776f bugfix in wrapper definition
464| * eb62aa2 implemented applyWrapper() now works on df.columns instead of df
465| * 293fc50 bugfix in wrapper definition
466| * b432bb1 intensityColumnsPerCondition in specificParams now contains the values from channelAliasesPerCondition in masterParams['schema'] !!!
467| * 6ae2975 bugfix in main(): global variables should be set in same loop as processDf calls
468| * d504f5d small bugfix
469| * 9aaf4e8 getPCA now omits NaN values
470| * b0accea updated getPCA and getHC calls to use allChannelAliases as intensityColumns parameter in getIntensities call
471| * 63e04ad implemented unnest() in main.py
472| * e479bec metadata is now multi-indexed
473| * 2349013 fixed bug in getProteinDF call
474| * 2de4e3a getIntensities now also accepts optional argument intensityColumns
475| * b1b3c8c adapted getProteinDF to handle multi-experiment data
476| * 1a7535d refactoring combineExperimentDFs(): no more fiddling around with intensityColumns or channelAliases: those names should be in-place from the beginning. BUT they are still called channelAliasesPerCondition and not intensityColumnsPerCondition in the masterParams !!!
477| * 0f4d36f implementing multi-experiment support in analyzeProcessingResults. implemented combineExperimentDFs(): all dfs are merged together with multi-index into allExperimentsDF
478| * 3da38ac implemented alias part of parseSchema
479| * c157cac implementing multi-experiment support in analyzeProcessingResults. all dfs are merged together with multi-index into allExperimentsDF
480| * a9f895d possibly simplified code in getProteinDF
481| * c14d098 implementing analyzeProcessingResult for multiple experiments
482| * d7bf91c adapted analyzeProcessingResult to work with dict as processingResults input type
483| * cbd82fa modifying input and output paths
484| * 0e32377 bugfixes
485| * 76e26fe small bugfixes
486| * fc40dec small bugfix
487| * 3fd6c9b implementing experiment names via masterParams['schema']
488| * dc2245d commented out parseSchema and refactoring getList -> parseExpression in getInput.py
489| * 8fdaded schema in masterConfig is now a dict which is written automaticaly by the web interface, containing the column aliases, intensityColumnsPerCondition and the config file location.
490| * 9c43b59 implemented multiple experiments: * refactoring: parameter files_in -> file_in * splits params in specificParams and masterParams * added schema.tsv and schema masterConfig parameter * added parseSchema to dataIO.py * intensityColumnsPerConditionPerexperiment is now a dict of dicts * added masterConfig.ini which contains also the analysis+report master parameters * each experiment now has its own config file * each processing step is inside a big for loop, going over each experiment.
491| * 8735977 added parameter identifyingNodes which now replaces masterPSMAlgo. Format: {"master": ["Mascot (A6)", "Ions Score"], "slaves": [["Sequest HT (A2)", "XCorr"]]}
492| * 5501df0 fixed small bug and import
493| * 82587b2 fixed imports
494| * f8c4432 moved getInput to separate file. main() now takes an extra argument configFilePath
495| * 2b4a89a implemented fixFixableFormatMistakes
496| * 182d307 implemented getDataFrame which now is superior to importDataFrame. made skeletons for fixFixableFormatMistakes and applyWrapper
497| * ec94159 Merge branch 'master' of ssh1.ulyssis.org:thesis
498| |\
499| | * f9c019d added more code and MA plot for comparing intensity and S/N SNR
500| | * 5cd8d40 commented adjustText import
501| * | 785252d makeHTML and HTMLtoPDF skeletons
502| * | 49e1658 diffMinFullProteins contains the differences between the min and full protein lists
503| * | 7864dc8 refactoring: getSortedDifferentials -> getSortedDifferentialProteinsDF
504| * | f489c79 getSortedDifferentials now only returns important columns (protein, significant, description, fold change, p-value
505| * | 0184c4f stuff
506| |/
507| * 91d3a0f fixed getSortedDifferentials
508| * bddeb8e added todo
509| * d5e5b3c cleaned up some code in getVolcanoPlot
510| * d3b2b46 combined annotation code in getVolcanoPlot
511| * fcb244c two volcano plots are produced: one for minProteinDF, one for fullProteinDF
512| * 3c3a034 added parameter labelVolcanoPlotAreas
513| * ee54949 volcano plot now has labels, if enabled through list of booleans labelPlot
514| * fcfbfc2 PCA plot now has labels (channel)
515| * d1fd19f do not save visualization objects to disk (yet)
516| * 9bf090f split getDataVisualization into getVolcanoPlot, getPCAPlot and getHCDendrogram
517| * 96c6f37 giving dendrogram labels
518| * 3a707c0 intensityColumns parameter is now obsolete in the config file (replaced by intensityColumnsPerCondition)
519| * 1bb9f93 getHC now first removes nan values
520| * 79bc9fa edited getHC because i misunderstood the instructions
521| * b57471b bugfixes
522| * 151ea97 added protein descriptions to proteinDF
523| * 571b6fe implemented getSortedDifferentials which sorts the protein dataframe according to fold change
524| * f1dbd5a removed bloat from main()
525| * 6ff848c enabled dendrogram and added some todos
526| * aefbb65 applySignificance docu
527| * 1417118 enabled and fixed bug in getHC
528| * 5e84dcf implemented volcano plot
529| * 126c117 added applySignificance() in analysis.py to create a column in the df which indicates the significance of the result
530| * baec5f0 implementing volcano plot
531| * d43c3b7 fixed PCA colouring
532| * be12d3d refactoring df -> normalizedDf where applicable (like analyzeProcessingResults)
533| * 6d4c764 PCA plot now colours channels according to their corresponding condition PCA plot now has distinguishable colors for arbitrary number of conditions.
534| * cfae79f stuff
535| * aa9bbbe bugfix in analysisResults dump file save operation
536| * c2559ac moved visualization save from analyseProcessingResults to generateReport
537| * bebb0f2 small bugfix
538| * 15ed138 moved dataVisualization BACK to report.py because we don't need to save figures anyway, and this seems like a morelogical place. It is no problem to separate the PCA and the plotting of its results. Just passing the PCA results and stuff inside Python (=inside memory) is better than saving figures to disk and reloading them. Plus, you can now separate the visualization from the calculations, which is nice for debugging.
539| * 715a100 bugfixes in dataVisualization
540| * 0035ded added comment with instructions on saving figures
541| * 089788f added minimum PCA_components check (>=2)
542| * 0105df3 implemented PCA plot in dataVisualization
543| * 4998e85 getPCA was incomplete
544| * 9812bdf dataVisualization started implementing PCA plot
545| * e531d5d refactoring: viz -> visualizationsDict
546| * 53ba72e implementing dataVisualization hierarchical clustering dendrogram
547| * 9e37af7 BUG in getHC: clustering causes segmentation fault SIGSEGV
548| * a5576b4 getPCA now imputess missing values by 1/#channels for its calculation (it does not apply the imputation to the dataframe!)
549| * e9e2332 bugfix in getPCA
550| * e2c87af moved global parameter definition from processing() to main()
551| * 6a1d7e4 removed obsolete code
552| * 4732bb8 implemented hierarchical clustering getHC()
553| * 6af4e16 added parameter PCA_components
554| * d38823e unknown
555| * e52f0f1 implemented PCA
556| * 5a1e2b6 changed PCA implementation: now using scikit-learn package
557| * d591aaa implementing PCA
558| * 75bd216 docu
559| * 6ce952c moved visualizations back to the analysis part (because you have to save figures and stuff BEFORE you generate the report)
560| * 613aa29 bugfix: processing/analysis results should only be loaded if the booleans doAnalysis/deReport are actually True...
561| * b943f2c added main.py:main() parameter doReport that can pick up on an earlier done analysis.
562| * b7d3f34 created report.py which is to do the visualization and report generation
563| * 08f2cdc implemented test function compareIntensitySN() to get an idea of the difference between the use of intensities or S/N values
564| * 89fa7dd small improvement + removed obsolete code
565| * 5a9f910 removed masked values from ttest output in applyDifferentialExpression
566| * b742339 small bugfix
567| * af13039 implemented applyFoldChange
568| * 6136883 skeleton improvements
569| * 493346c small bugfix
570| * 1cb80af dummy implementation of applyFoldChange
571| * bfa46dc made results output nicer (changed column order, put protein back into the columns after the analysis is done)
572| * be3c9d5 in proteinDF the lists are now actually lists instead of series only the pvalues and adjusted pvalues are saved to file
573| * b8c65e3 bugfix: omit NA values in t-test
574| * b04f50e small bugfixes
575| * cdf0e79 small bugfix
576| * 1237552 bugfix in data exports (3rd return variable must be a dataframe)
577| * 64d8777 small bugfix in data exports (no delim specified)
578| * 47a0082 small bugfix
579| * d1a51df bugfix in applyDifferentialExpression assignment
580| * c6c3c17 small bugfix
581| * b713d07 bugfixes in appliDifferentialExpression bugfix in getProteinDF: now is a 3Nx1 array per condition instead of a Nx3 array
582| * cd80228 bugfix: manual cast from int64index to list. very ugly
583| * 45e33aa bugfixes in getProteinPeptidesDicts
584| * 43ecf02 bugfixes in getProteinPeptidesDicts now warns when peptides without master protein accessions detected (numGroups == 0)
585| * e64078b bugfix getProteinPeptidesDicts numGroups is of type int
586| * cb5179d refactogin proteinDF() -> getProteinDF()
587| * 409d869 removed illegal filename test case from files_path because it was annoying to debug.
588| * 9c511b8 import spelling mistake
589| * f6b5b53 abspath->relpath
590| * 3426c29 exception clarification
591| * ac19bf2 bug UNFIXED: this_proteinDF is empty when it enters applyDifferentialExpression
592| * e647cae bugfix in applyDifferentialExpression
593| * e5ee2fa small bugfix in applyDifferentialExpression: .loc[] -> [] for new column definitions on dataframes
594| * efb9915 bugfix in creation of empty dataframe
595| * 291708d small bugfixes
596| * c5f1aa1 bugfixes (mostly apostrophes "" around integers in config file)
597| * b63a34e bugfix for multiple input files which is not completely implemented in main()
598| * ca03332 bugfix for multiple input files
599| * 85c64e4 config.ini update + bugfix
600| * 7c2c571 bugfix
601| * 6c42b9d bugfix in json module import
602| * 62265af made data analysis callable without processing using a pickle python object written at end of (previous) data processing step
603| * ccd987b fixed bugs: refactoring artifacts
604| * ad3c964 post-merge chaos
605| |\
606| | * f92de2a added booleans doProcessing and doAnalysis to control which parts of the workflow to execute;
607| | * 1c86f4f implementing the workflow for allowing multiple experiments
608| * | c289efb Merge remote-tracking branch 'origin/master'
609| |\ \
610| | * | b34b128 started elaborating analysis workflow
611| * | | 10453c5 bugfix: didnt add parameters everywhere in dataIO
612| * | | 61c30f5 removed: dataIO.py:parseList() now json.loads handles lists, can also handle nested lists
613| * | | 08d3561 added parameters alpha, FCThreshold, intensityColumnsPerCondition, pept2protCombinationMethod (mean/median)
614| * | | 2463d73 bugfix: indexing in getProteinPeptidesDicts
615| * | | 959a5d5 fiddling around with imports
616| * | | d188d41 comments
617| * | | 0b849fd refactoring: foldChange -> applyFoldChange
618| * | | 34ad786 refactoring: differentialExpression -> applyDifferentialExpression jsut modified the input df
619| * | | b52cb33 refactoring: shadowing variables
620| * | | 087ebc0 removed obsolete code
621| * | | 0fd38dd refactoring: proteinIntensitiesPerConditionDF -> proteinDF implemented proteinDF
622| * | | 922368e refactoring: setGlobals -> setProcessingGlobals
623| * | | 529cff5 implemented differentialExpression
624| * | | d047ef4 started elaborating analysis workflow skeleton is as good as finished
625| |/ /
626| * | 9f64477 setIntensities now handles both ndarrays and dicts
627| |/
628| * d4d8bbd refactoring dataprep.py -> dataproc.py and very small bugfix
629| * 326c943 refactoring: isotopicCorrectionsMatrix -> isotopicCorrection_matrix
630| * 78a02b4 refactoring: remove_ExtraColumnsToSave -> removalColumnsToSave
631| * f1b1fc4 bugfix in setMasterProteinDescriptions: can now handle case where there is no Protein Descriptions colmun present (just does nothing but throw a warning in that case).
632| * 5a951d5 implemented analysis.py:getProteinPeptidesDicts() (including documentation)
633| * b99032b refactoring: requiredColumns -> wantedColumns better reflects the nature of the list (see previous commit changes for context)
634| * 970697f refactoring and re-implementation: getRequiredColumns -> removeObsoleteColumns because it now removes all except the wanted ones, instead of just selecting the wanted ones. This is because all the required columns may not be present but may still be non-essential (Xcorr, Ions score!) in other words they are not in noMissingValuesColumns
635| * b64c3ab getBestIndices now still works if there is not PSM slave score column
636| * cbf8941 Merge branch 'master' of ssh1.ulyssis.org:thesis
637| |\
638| | * e97d3d1 Merge remote-tracking branch 'origin/master'
639| | |\
640| | * | 99e985c small bugfix
641| * | | d80d92f changedconfig file tonot use Kurt's data
642| | |/
643| |/|
644| * | 39883d5 bugfix in isotopicCorrection: didnt properly return index of rows with NaN values
645| |/
646| * 7596222 implemented warning and solutiong (pick first index in list of duplicates) for when there is no BestIndex for a list of duplicates.
647| * 6bb0052 removed obsolete RT flagging code inside combineDetections()
648| * d4dfa5c added getNoIsotopicCorrection to analysis.py to record which detections received no isotopic correction.
649| * 6d8ca7d added duplicates column to getRTIsolationInfo
650| * 63a6219 added metadata in main.py and getRTIsolationInfo in analysis.py
651| * 5dba56d added date parameter
652| * 629ddff bugfix in combineDections: now returns a dict of intensities instead of just one vector with 6 values.
653| * ff6902e combineDetections for geometricMedian: rows should be normalized before applying geometric median, otehrwise the norm isnt conserved to one.
654| * cb22aaa bugfix in getIntensities for cases where df contained only one entry (then becomes a series)
655| * 94b8375 geometricMedian now throws away detections with missing values; 'mean' method now employs nanmean isntead of regular.
656| * 75599a4 bugfix; removed intensityColumns dependency and used new version of getIntensities (see prev commit) instead
657| * 04207ba getIntensities can now handle a selection of indices
658| * 99fb008 no collapsePTM allowed
659| * bb836ab bugfix + clarification in setMasterProteinDescriptions
660| * 179735c implemented setMasterProteinDescriptions
661| * 007464d comments
662| * 2a79782 bugfix
663| * 163e566 bugfix
664| * eda0579 refactoring: unusedDetectionsInOneFile -> removedDataInOneFile_bool
665| * f6e4c5d bugfix in dataprep.py:setGlobals
666| * fc0ebda added functionality to save removedData as one file
667| * f053cc1 simplified code redundancy in undoublePSMAlgo and added unusedDectionsInOneFile_bool parameter
668| * 3b62e6e added noMissingValuesColumns and remove_Extracolumns parameters so that very few columnsToSave entries are hardcoded (only those who will never change)
669| * c33b4dc removed collapseRT_bool
670| * 67074fa fixed bug in getDuplicates for collapse PTM: annotated sequence is now first .str.upper()-ed before groupBy annotated sequence.
671| * 010d1e9 fixed off-by-one error in getRepresentativesDf --> BIG BUGFIX off-by-one + race condition during data manipulation = results non-deterministic
672| * 8d59ddf fixed off-by-one error in getRepresentativesDf --> BIG BUGFIX off-by-one + race condition during data manipulation = results non-deterministic
673| * 06a5188 rewrote getIntenseIndices stuff to also use a dict
674| * e4325dd rewrote getIntenseIndices --> getIntenseIndicesDict (self-explanatory) { bestIndices : intenseIndices }
675| * ff7f879 some refactoring even more df[ --> df.loc[ and refactoring to avoid shadowing and rewriting collapse() so that bestIndices --> bestIndicesDict is properly used
676| * ad17ac3 refactoring even more df[ --> df.loc[
677| * 3e148d7 refactored getBestIndices (code redundancy)
678| * 9d07f46 refactoring to avoid variable name shadowing
679| * 23e320c seems to be a VERY WEIRD bugfix ( you know, the bug where the output of the program was randomized and suddenly the dataframe had itself inside of itself on some index and on the other indices just 22 times "somestring" while it looked completely normal on inspection through the debugger but not when you called its indices... ) where I now changed representativesDf['Degeneracy'] ------> representativesDf.loc['Degeneracy'] and it seems to be solved
680| * 68f2dc4 small possible bugfix in getRepresentativesDf
681| * 09a30ec assert statement
682| * 8783559 simplified removedData construction (direct from df instead of list of dataframes per duplicates group)
683| * d91dfd2 documented that the order of adding representatives in getrepresentativesDf is important.
684| * a4cb99b bugfix in groupByIdenticalProperties: now also checks if the candidate duplicates group consists ofmore than 1 member AFTER all groupby's have been done. Also, now uses filter() ahead of the for-loop instead of an if statement inside the for-loop.
685| * c5339aa removed obsolete code that replaces groupByIdenticalProperties
686| * 85639c1 small performance increase in groupByIdenticalProperties
687| * cc22e1c badConfidence->confidence (removedData name)
688| * 538af1f fixed performance issue involving chained assignment
689| * a166ddd isotopicCorrections now doesnt try to correct detections that have NaN values in their intensities. Gives a warning if this occurs at least once
690| * b4d2a51 removed profiler inside code, using profiler.py
691| * 1c63bba added profiler
692| * 14ee2ad bugfixes in undoublePSMAlgo
693| * 377f1da bugfixes in undoublePSMAlgo
694| * 2f02feb removedData in collapse: now saves representative first scan instead of index (because index is unknown to the user)
695| * 61280e4 removed some SettingWithCopyWarning
696| * ea1a8af exportData can now handle dataframes in removedData object and write them separately to dsik.
697| * 2bfdd5a bugfix in: unified all removedData formats. collapse function's removedData is now one dataframe containing all collapsed detections along with their parent ID.
698| * 45c575b small bug in undoublePSMAlgo assertion
699| * a4123b2 unified all removedData formats. collapse function's removedData is now one dataframe containing allcollapsed detections along with their parent ID.
700| * 7ea31b7 removedData in undoublePSMAlgo is no longer a tuple (master PSMAlgo clear from removedData PSM score column name
701| * 73bc79b removedData in undoublePSMAlgo is no longer a tuple (master PSMAlgo clear from removedData PSM score column name
702| * 4b8344d bugfix in sanity check
703| * 490ab6a bugfix in getRepresentativesDf
704| * a7a2926 bugfix in getRepresentativesDf
705| * 25c019f main() now takes testing and writeToDisk as argument booleans. implemented testDataComplementarity
706| * d81d329 documentation
707| * 3b94039 bugfix in exportData
708| * 211f50c exportData documentation
709| * 2d04940 config filename extension removal
710| * f1945ec elaborated exportData for objects (like removedData)
711| * 28709e8 bugfixes; getBestIndices
712| * 7e920f8 bugfixes; getrepresentativesDf
713| * 0af69f6 bugfixes; groupByIdenticalProperties
714| * 724c62c bugfixes; made toDelete in undobulePSMAlgo into Set
715| * 8dd5d9f documented getRepresentativesDf fixed indentation bug and renamed indices->"group of indices" in duplicateLists explanation
716| * caefba5 documented getIntenseIndices
717| * 0e7632b updated getDuplicates:groupBy... documentation and documented getBestIndices
718| * 4d7c8fd implemented degeneracy propagation in getRepresentativesDf()
719| * cf58359 refactoring: scanDict -> byFirstScanDict
720| * 4241303 elabroated combineDetections(), removed getNewIntensities()
721| * cd73089 added locations for flagging and implemented flagMaxReporterVariance in combineDetection() in collapse() in collapse.py
722| * c644aa5 cleaned up some old code in getRepresentativesDf
723| * 628b203 implemented the saving of the removedData in collapse.py:collapse()
724| * 1b4da30 implemented getRepresentativesDf()
725| * acdd417 BIG CHANGES: rewriting collapse() and all its functions. representatives are now a COPY of the best matching of the duplicates, and their intensities are modified if method is not bestMatch.
726| * 86ce581 implemented getDuplicates() to use the groupby function. Also made it compacter using grouByIdenticalProperties(), but also left old code intact and implemented switch youreFeelingLucky boolean
727| * f74d462 reworked undoublePSMAlgo to use the groupby function
728| * bf603d3 added masterPSMAlgo to collapse.py:getRepresentative()
729| * e61ce91 removeBadConfidence documentation
730| * 6dea8d4 added degeneracy column when collapsing
731| * d82ea86 refactoring: columnsToSave -> collapseColumnsToSave for collapse.py usage everywhere
732| * 70b39d6 fixing some bugs in possible collapse function combinations
733| * d437ade added and implemented removeMissing in dataprep.py
734| * 176cfad added and implemented removeMissing in dataprep.py
735| * f024350 refactoring: made columnsToSave a parameter in config.ini
736| * 3ad66c5 unified colsToSave for all collapse cases
737| * c952437 fixed bug in configparser for # symbols holy shit why do people put # signs in their legit output
738| * f591b37 added setIntensityColumns in dataprep.py elaborated selectEssentialColumns in dataprep.py included essentialColumns and intensityColumns as parameters in config.ini added parseList in dataIO.py
739| * 9e3792e added removeBadConfidence function along with its two parameters
740| * 972f8a8 fixed some import bugs due to refactoring
741| * 28591cd refactoring: moved all collapse-related functions to collapse.py from dataprep.py and modified their documentation added selectEssentials()
742| * b4cfc4c refactoring: collapsePSMAlgo -> undoblePSMAlgo
743| * b8b044c fixed bug in collapse function calls in main()
744| * ec5b77c refactoring: collapsePSMAlgo_master -> masterPSMAlgo
745| * 1ed6f00 refactoring: generalized all collapse functions into one collapse function (except collapsePSMAlgo because its not a real collapse) added function getRepresentative, made combineDetections more explicit
746| * 9b66d07 refactoring: generalized all collapse types/methods and centermeasures into one collapse_method parameter, used by all collapse functions.
747| * 2a7c788 refactoring: created getNewIntensities() function which replaces the identically named functions in the body of the collapseXXX() functions.
748| * d1fd418 refactoring: created collapse() function which now handles almost all of the commands (NOT function definitions) in the collapseXXX() functions.
749| * 359336c implementing new dataprep.py:collapsePSM()
750| * 6c345d3 implemented sanity check: main() checks whether there are still identical (!= duplicate) peptides after collapsePSMAlgo()
751| * 8112234 implemented sanity check: main() checks whether there are still duplicate sequences after all collapses have been applied
752| * 28e860b implemented sanity check: checkTrueDuplicates of collapseCharge() now asserts that you ran collapseRT() implemented sanity check (test): also assumes you ran collapsePSMAlgo()
753| * 40ec914 decided on collapseRT method and centerMeasure structure. Implemented barebones version in collapseRT() and created new fucntion combineDetections() which contains the centerMeasure part.
754| * 7dc2990 decided on collapseRT method and centerMeasure structure. Implemented barebones version in collapseRT().
755| * 9aea745 refactoring: collapseRT_centerMeasure_addition -> collapseRT_centerMeasure
756| * a4bfb31 created MS2IntensityDoesntMatter() dev method
757| * 6465bb0 refactor collapseRT_centerMeasure_intensities -> collapseRT_method
758| * a9e115c refactor collapseRT_centerMeasure_reporters -> collapseRT_centerMeasure_addition
759| * aed17e9 added missing removedData assignments in main()
760| * c15e340 added collapseCharge_bool dependency on collapseRT_bool in getInput()
761| * db593a2 refactor: .iloc->.loc for fear that positional instead of label-based indexing might fck things up, although it looks like it doesn't.
762| * b3a75e3 tested isotopicCorrections(): it works
763| * eb0aa6b bugfix: isotope impurities should be imported as float64 instead of int64
764| * b649bc8 refactoring: channels->reporters
765| * 21f478a bugfix in getIsotopicCorrectionsMatrix()
766| * c45df6e bugfix in getNewIntensities()
767| * 897c0d7 isotopic corrections documentation
768| * 3f3d4b2 stuff
769| * a811c84 implemented dataprep.py:isotopicCorrections()
770| * 88b7cff added ICM check
771| * 7e452b7 moved isotopiccorrectionsmatrix to a separate file.
772| * c5edfd0 bugfix in removedData
773| * 65b1931 implementation collapseRT() and refactoring maxRelativeReporterVariance
774| * 4aef562 moved getDuplicates() OUTSIDE the collapseCharge() function so that it can also be used by the collapseRT() function. It now also needs the dataframe and a checkTrueDuplicates() function.
775| * 95cf115 modified the way collapseCharge retains deleted data: associated each duplicate row with its matching firstOccurrence using a dict
776| * b1c9e7e documentation minor name change
777| * 7f86f14 name change updateFirstOccurrences() -> getNewIntensities() documentation
778| * 9922af6 modified updateFirstOccurrences() and collapseCharge() to produce a dict for setIntensities() input
779| * d871821 implemented setIntensities using a dict as input
780| * ab98e23 implementing collapseCharge() but this is a freaking b*tch there are still issues with a) how to detect peaks that ought to be collapsed and b) how to combine the intensities of the duplicates
781| * ef62c06 implementing collapseCharge() but this is a freaking b*tch
782| * 03d38bf added ext2delim and delim2ext functions. streamlined importDataFrame() using those functions
783| * 6225297 added ext2delim and delim2ext functions. streamlined importDataFrame() using those functions
784| * fec450e fixed bugs in dataprep.py:collapsePSMAlgo() changed dataIO.py to use the config.get-methods instead of castin, because the getboolean() parser is very neat and bool() casting is dumb.
785| * bf93820 fixed bugs in dataIO.py:getInput()
786| * 7591664 fixed bugs in dataIO.py:getInput() but still need to fix many (apply variable changes to params dict)
787| * fbefea4 dataIO.py:getInput() now uses a config parser created config file config.ini
788| * f55be14 Merge branch 'master' of ssh1.ulyssis.org:thesis
789| |\
790| | * 7bd23c8 implemented collapsePSMAlgo() (still buggy)
791| * | 0dc8c37 fixed a refactoring error
792| * | 1623c16 removed some whitespace
793| * | 46c2218 tweaked performanceTest to not use dataIO.py
794| * | 578c5df added isotope impurities matlab scripts
795| * | 47aa2c3 cleaned up as much testing stuff from main method as possible
796| * | 0b98aca tested for isotope impurities correction by constand or not (NO)
797| * | 40b4297 added dataType param to exportData
798| * | 0787eb6 implemented collapsePSMAlgo() (still buggy)
799| |/
800| * ca8775f added test_dataprep.py (barebones)
801| * a355b30 added test_dataprep.py (barebones)
802| * 312e917 isotopicCorrection: do not output corrected intensities to user.
803| * fc74323 implemented removeIsolationInterference() and removed addColumns in favor of a dict removedData.
804| * af6724b stuff
805| * 26ccbf8 skeleton of peptide to protein map in analysis.py:mapPeptToProt()
806| * 79f68aa a LOT of bugfixes in the skeleton / parameter chaos
807| * 3c29c64 forgot isolationinterference parameters in dataIO.py
808| * 98ea76d small bugfix
809| * de4d02a for now I think all the necessary parameters are included
810| * daf8104 added DEFoldThreshold
811| * ac4ed19 added collapseRT_maxRelativeChannelVariance parameter (damn these names are getting long.....)
812| * 92e4637 addid removeIsolationInterference (barebones)
813| * 8e9472e modified exportData to handle both a path plus a filename instead of only a full filepath
814| * e0dcc6c removed the importData() function (importDataFrame() does all that is needed).
815| * 9c80073 further explicited workflow skeleton and function skeletons (added parameters). added function (barebones) to add columns according to workflow format for retaining to-be-deleted info. also added more parameters checks as parameters are added.
816| * 48bdb19 renamed selectIntensities->getIntesities moved getIntensities to dataprep (from dataIO) added setIntensities (barebones)
817| * 80e99fe stuff
818| * 1ddc7b5 dataIO.py:getInput() now returns the parameters inside a dict "params"
819| * 11ae689 built the entire workflow skeleton inside the main.py:main() and made barebones files (analysis.py, dataprep.py) with the required functions.
820| * 5ab7693 added barebones test file test_dataIO.py
821| * b3eea3e fixed basic unit tests for test_constand
822| * 349e226 added header_in argument to importData and importDataFrame in dataIO.py
823| * e1c5a74 made basic test_constand
824| * 610a344 removed nan values in returned convergenceTrail
825| * 2124638 added checks to dataIO.py and small refactoring
826| * 9d7d50b added checks to dataIO.py and small refactoring
827| * 3d827c5 stuff
828| * 24e2d95 docu of constand.py completed (for now)
829| * a4942ad docu of dataIO.py completed (for now)
830| * 19ee041 started documenting (dataIO.py)
831| * 4934496 initialized test file for constand.py:constand() created dataIO.py and migrated code from main.py into it
832| * 319f45b added exportData function
833| * 89f6750 added exportData function
834| * 96cf907 mail.py:selectIntesities is working now
835| * 7b968c4 fixed bug in importDataFrame elaborated importDataFrame (added more security checks) elaborated getIntensities (not finished yet)
836| * eaef437 added main.py:selectIntensities (empty for now) which should extract the intensities from the dataFrame
837| * b2c9601 elaborated profiler function (though some more tweaking needed + need to add line_profiler switch, now only has cProfiler)
838| * ebe90a9 elaborated importData function to contain test stuff (does not call importDataFrame anymore)
839| * e01b9f7 added profiler.py for testing the performance
840| * 20b7d10 elaborated main.py:importData and renamed to importDataFrame. importData should become a different function.
841| * a97c9d3 (moved from constand_matrix_mult) aesthetics + main.py:importDataFrame->importData
842| * 88b8ffa aesthetics
843| * 4cc3dd2 added performanceTest method so that i dont have to copy paste that code each time.
844| * 4b71e94 bug in constand.py:convergencetrail gefixt; NaN waarden en overeenstemming met MATLAB output succesvol getest.
845| * ec5579d import_MB....m script elaborated (is eigenlijk gebeurd maar is toen niet bij in de commit geraakt precies)
846| * 90a6abe bugs in constand.py gefixt. Algoritme werkt nu met Ri en Si per iteratie ipv enkel het totale product R en S. (maar R en S worden wel gaandeweg berekend zodat het niet nodig is extra memory te voorzien)
847| * c7e9087 implemented constand.py but it still has bugs; elaborated the main skeleton (getInput etc)
848| * f461257 more matlab test scripts, created constand.py
849| * 4e4c54b main file skeleton
850| * 4ad11ae initial commit
851* 927bf60 PACKAGE INIT RESTRUCTURE: created setup.py and moved all files into the parent dir via a trick (create subdir first, move the whole tree into it and then remove the original folder from the path after copying the subtree to the parent dir as done in http://stackoverflow.com/questions/3212485/how-do-i-re-root-a-git-repo-to-a-parent-folder-while-preserving-history)
852* 8b9bc30 moved unnest() from dataIO.py to new file tools.py
853* 8308650 MAIN SCRIPTS CLEANUP: removed devStuff()
854* f74caff MAIN SCRIPTS CLEANUP: moved, docu and refactoring of dataSuitabilityMA->testDataSuitability to scripts dir
855* 95a76bb MAIN SCRIPTS CLEANUP: moved, docu of compareDEAresults to scripts dir
856* 2d4ca34 MAIN SCRIPTS CLEANUP: moved, docu of intraInterMAPlots to scripts dir
857* 3f6970c MAIN SCRIPTS CLEANUP: removed compareAbundancesIntSN()
858* 22351e2 MAIN SCRIPTS CLEANUP: moved, docu and refactoring abundancesPCAHCD()->PDAbundancesPCAHCD to scripts dir
859* f871779 MAIN SCRIPTS CLEANUP: moved, docu compareICmethods() to scripts dir
860* 12c3408 MAIN SCRIPTS CLEANUP: moved, docu and refactoring: compareIntensitySN->testIntensitySNDifference to scripts dir
861* 79a780b moved fontsize and fontweight global variables from main.py to __init__.py
862* 53bc8a0 MAIN SCRIPTS CLEANUP: moved MA(), scatterplot(), MAPlot(), boxPlot(), RDHPlot() to tools.py in the scripts dir
863* 55375d0 MAIN SCRIPTS CLEANUP: removed testDataComplementarity()
864* 745fc8e MAIN SCRIPTS CLEANUP: removed MS2IntensityDoesntMatter()
865* 0cb749b MAIN SCRIPTS CLEANUP: moved and refactoring: isotopicCorrectionsTest()->testPDIsotopicCorrectionsEffect.py
866* 00c8887 MAIN SCRIPTS CLEANUP: moved, docu and refactoring: isotopicCorrectionsTest()->testPDIsotopicCorrectionsEffect.py
867* df057a9 MAIN SCRIPTS CLEANUP: moved and refactoring: isotopicImpuritiesTest()->testIsotopicCorrectionNecessity.py
868* 0ffdaab MAIN SCRIPTS CLEANUP: removed performaceTest()
869* 2a6b3f3 moved profiler.py from constandpp to scripts repo
870* 29caf77 applyFoldChange() docu clarification
871* 929f115 smaller scatterplot fig size
872* 3b69543 fixed bug in scatterplot()
873* 51a772e added int64 as viable data type for constand
874* 348812e added default parameters to constand function
875* c62e3e8 fixed small bug in scatterplot. Probably the MAPlot function won't work that properly anymore.
876* 59ec2b4 added scatterplot script function to main.py that MAPlot now depends upon
877* e104e8b changed emailadress to UHasselt address
878* c0450e3 added __init__.py and moved credentials there
879* b7c3eff fixed bug due to renaming
880* ba9336c fixed unused imports
881* 8073299 fixed webFlow unused import
882* 13dbc5e docu send_mail in web.py
883* ecacd2f deleted obsolete constructMasterConfigContents in web.py
884* 313cad9 line endings bullshit
885* b348b97 refactoring modified TMT_ICM->TMT_IDT files in job folder stuff
886* 915bd44 refactoring modified TMT_ICM->TMT_IDT files in job folder stuff
887* 35ec345 docu startJob in web.py
888* 8d272f5 stuff
889* 366bea3 fixed docu mistakes in setJobCompleted and setJobFailed
890* 976a1c4 docu DB_setJobCompleted in web.py and rename jobDirName->jobID
891* 973a531 docu DB_setJobReportRelPaths in web.py and rename jobDirName->jobID
892* 031ec37 docu makeJobConfigFile in web.py
893* 6bcd574 docu updateWrappers in web.py
894* 698a852 added todo because there is a file manipulation function location inconsistency between dataIO.py and web.py
895* f0b3200 added jobs symlink in static folder to git ignore
896* 7834be3 docu updateConfigs
897* 641791b docu DB_getJobVar in web.py
898* e0a5b08 docu DB_checkJobExists and DB_insertJob in web.py
899* a1570b6 docu updateSchema in web.py
900* 5152f2d docu saveFileStorage in web.py
901* e08514b docu hackExperimentNamesIntoForm
902* 55b4312 docu web.py file description
903* 9cd43f7 docu newJobDir in web.py
904* fcee300 webFlow.py docu file description
905* a5d891e docu hackImagePathToSymlinkInStaticDir
906* d1083bc created injectColumnWidthHTML() in report.py to fix the relative column widths in the pandas-generated differentials table
907* 8e66c80 used the pandas to_html() function to generate differential proteins table HTML instead of trying to loop over a dataframe in Jinja
908* 2c300d2 stuff
909* 5382908 fixed obsolete argument bug in getPCA call to getMarkers
910* 1c07287 added job failure message to jobInfo if job is done but failed
911* 28798f4 put try-except around mailer.send(msg) in send_mail() in web.py and used a flash message to inform the user via the web. So now the application des not stop if a mail doesnt get sent
912* cfb1815 removed obsolete code in views.py
913* 3ecdf29 fixed bug in jobInfo() in views.py if the ID was invalid. Updated the error message to contain link back to /jobinfo
914* ae7d9d6 docu jobInfo in views.py
915* cfdc53b moved allJobsDir and jobDB parameters from __init__.py jo config.py and made uppercase
916* 8ac254e updated config.py
917* d061c52 added __pycache__ to gitignore
918* 3a37ef5 editing: jobInfo docu in views.py and parameters in config.py
919* fd75905 docu jobSettings in views.py
920* f40d8c4 docu up to jobSettingsForm in views.py
921* 2566a01 docu __init__.py about the views import statement at the end of the file
922* 27fc597 docu forms.py
923* 0fa2823 docu config.py
924* 1bb5443 docu __init__.py and rename of parameter DB->jobDB
925* 0f3e073 docu collapse() in collapse.py and removed obsolete code
926* 016cf04 docu getRepresentativesDf in collapse.py
927* 10dced3 docu getIntenseIndicesDict in collapse.py
928* 85a0648 docu getBestIndicesDict in collapse.py
929* 9e327d4 docu combineDetections in collapse.py
930* 3690908 docu groupByIdenticalProperties
931* 0185e20 refactoring: getData->importExperimentData
932* 341e2fb refactoring: getWrapper->importWrapper
933* 7f3f1d6 refactoring: getIsotopicCorrectionsMatrix->importIsotopicCorrectionsMatrix
934* f636069 added .gitignore file after git-extras package installation
935* 4313d44 docu dataIO
936* cb20008 main.py fixed some layout and comment details
937* a39f672 commented out writeConfig() + docu (partial) dataio.py
938* 629e3ba refactoring masterConfig->jobConfig and docu getinput.py
939* 2f79c32 docu runweb.py
940* 17e307d docu makeHTML in report.py
941* 9f662ab docu HTMLtoPDF in report.py
942* 925117d docu getHCDendrogram in report.py
943* be5b471 docu and removed obsolete code getPCAPlot in report.py
944* 740f99a docu getVolcanoPlot
945* f6bfe7e docu and removed obsolete argument getMarkers() in report.py
946* 6d5c13e docu getColours
947* 939e165 convert indents to tabs
948* 29658b2 docu getColours
949* 4e1de9e docu generateReport in reportFlow.py
950* c7e4b1b fixed main.py imports
951* 77e690a (origin/webinterface, webinterface) MERGE UPDATES FROM BROKEN TO ORIGINAL REPO started docu reportFlow.py
952* 4525abd MERGE UPDATES FROM BROKEN TO ORIGINAL REPO fixed imports in main.py
953* 38d1e3c MERGE UPDATES FROM BROKEN TO ORIGINAL REPO docu processing.py
954* d48fe5b MERGE UPDATES FROM BROKEN TO ORIGINAL REPO docu processingFlow.py
955* 7fcf429 MERGE UPDATES FROM BROKEN TO ORIGINAL REPO docu analysis.py
956* 1f0d796 MERGE UPDATES FROM BROKEN TO ORIGINAL REPO refactor applyDifferentialExpression->testDifferentialExpression docu testDifferentialExpression
957* e8b2b89 put all paths in app config
958* 8078d9a Merge branch 'webinterface' of ssh1.ulyssis.org:thesis into webinterface
959|\
960| * 26ab051 moved updateConfigs to web.py; fixed some bugs
961* | f289d70 set up mail settings
962* | 666c0c6 setup mail attachments
963* | 438531e added email to forms
964* | 6370b7c set autorefresh to 5s
965* | a94be62 fixed the fucking form bugs satan will have my soul for this dirty code
966* | 34a34da fixed bug url_for image generation bug by just FUCKING NOT USING IT (nad not setting SERVER_NAME because it breaks everything)
967* | 0b7191d bug in url_for usage in report.html when SERVER_NAME is not set
968* | bece656 bugfix: getFile() view should use request args, not url argument
969* | 29955ad weird bug in analysisFlow for a missing parameter that should have surfaced 1 month ago
970* | 72fe1ea bugfixes in volcanoFullPath declaration (now always defined, sometimes as None)
971* | c3dff97 implemented makeHTML() and HTMLtoPDF
972* | 2330d3d added some metadata
973* | 4032dd5 exporting figures now returns their full path
974* | 2109fa0 made report template report.html
975* | 3decbd4 HTML and PDF reports now show via jobInfo()
976* | 7a9bc66 fixed jobInfo refresh, now working on showing the HTML and PDF reports
977* | 57a08a8 FINALLY fixed bug where the program didn't run
978* | 463b3a7 fixed report generation flow skeleton
979* | e5ddb19 startJob now returns Popen object
980* | 1448ad6 bugfixes
981* | 66e1ae0 added delim_out to jobSettings form, fixed app context bug (?)
982* | e2505f2 fixed some busg (mostly in startJob)
983* | 53b3cdd fixed isDone in jobInfo
984* | 9597ca3 removed DB_close calls
985* | 68d00b9 added empty isotopicCorrection_matrix entry in updateSchema
986* | 4d4b51c fixed BASEPROCESSINGCONFIG file path in config
987* | e346258 modified startJob to set stdin,stdout,stderr=None
988* | 33c8b42 implemented startJob
989* | 2d46d95 completed webFlow transfer to jobSettings() in views.py; implemented DB_setJobReportRelPaths, DB_setJobCompleted, DB_setJobFailed
990* | 74d9e9c refactoring in webFlow.py: master...->job...
991* | b2dff8a moved updateConfigs and updateWrappers to web.py
992|/
993* 71972b8 implemented DB_checkJobExist(), DB_close(), DB_insertJob(), getJobVar() in web.py
994* 66558ac fixed updateSchema; implementing what should be done after updateSchema
995* 1392f51 FINALLY fixed experiment schema labels :D by adding hackExperimentNamesIntoForm
996* e4c3c0a still fixing experiment schema labels
997* 377d81b filled jobSettingsForm with more options
998* 04f2670 implemented jobInfo with getHtmlReport and getPdfReport
999* 2a83bcb added jobstartedMail
1000* ec2875a included html save procedure in dataIO
1001* 34bba30 added sqlite3 database use
1002* de9caa3 implementing jobSettings forms
1003* e6dec51 fixed schema upload added newjob_content.html to be included in newjob.html and index.html
1004* cbd0afc added config file
1005* 3a86333 created jobSettings(), which gets called after successful schema upload.
1006* 74f332e renamed schemaForm to newJobForm moved newJobDir() from webFlow.py to web.py implementing schema file check after upload
1007* 4f80d6d made schemaForm()
1008* 46a69c6 added newjob.html and included it in index.html also added forms.py
1009* 81f938b added base.html and documentation.html and now uses jinja block extentions
1010* dd3958b added css added completed.html
1011* ea98a14 implemented mailer (untested) implementing css
1012* a53310c implemented getFile for rendering images (or files) from different locations
1013* bcdaba2 created link to dummy report file
1014* 04a0549 disabled plot show because it disallows the program to end without closing the figures
1015* ac15205 moved web.py and webFlow.py to web folder and fixed imports
1016* 0c0c360 created web subdir with views.py
1017* 1810d6b removed tests
1018* d50b1dd Merge branch 'master' of ssh1.ulyssis.org:thesis
1019|\
1020| * 2daba78 removed m files
1021* | 511e043 removed m files and other obsolete stuff like tests
1022|/
1023* fc7c409 adapted runweb.py to be open for any host
1024* ca41a68 added runweb.py to run a web interface
1025* 0f2d77e split off constructMasterConfigContents and TMT2ICM from webFlow.py into web.py
1026* 4928959 volcano plot limits
1027* bfe0fbe volcano plot limits
1028* 556da73 stuff remaking plots
1029* 1c3b319 PCA and volcano plot label size is 20 and marker size is now 80 and 160 in legend
1030* 12d1348 compareICmethods() improved
1031* 7c198ba MA plot improvement font is still normal but can be adjusted through fontweight
1032* a8e0a96 MA plot improvement
1033* 9b3638e adapted siotopicCorrectionsTest
1034* 1505655 schema is now an OrderedDict, so that its items() order is decided by the order in which the keys were added, i.e. the order of the experiments in the schema file uploaded by the user. this is the best fix since sliced bread
1035* 7096876 PCA_components parameter is now hardcoded in the webflow to be 2
1036* 3f76be8 comments
1037* edf09f9 removedData are now saved in a separate removedData folder per processing job
1038* 471a83d implemented numDifferentials parameter to control the amount of proteins in the DEA list of the report
1039* 8f6c1d0 Merge branch 'master' of ssh0.ulyssis.org:thesis
1040|\
1041| * ab05931 undoublePSMAlgo now checks whether there actually are enough PSMALgo's to undouble, even if undoublePSMAlgo_bool = true
1042* | 2c85b83 refactoring: minProteinDF_bool -> minExpression_bool; fullProteinDF_bool -> fullExpression_bool
1043* | 2fa3749 adapted analysis workflow to only calculate minProteinDF and fullProteinDF if specified by minProteinDF_bool and minProteinDF_bool adapted reportFlow as well
1044* | 3c1b487 removed todo
1045* | 8b49e8c undoublePSMAlgo now checks whether there actually are enough PSMALgo's to undouble, even if undoublePSMAlgo_bool = true
1046|/
1047* 3693bcd changed ICM string in TMT style isotopic corrections table to IDT string
1048* d4f2548 fontsize 30
1049* 4b1888b improced RDHplot to take a quantity argument specifying fidderence between what quantity
1050* 1269001 MAX_abundances_noTAM hard code in compareDEAresults()
1051* b295b6b updated fontsize lines (now refer to variable fontsize)
1052* bb3f0d9 implemented dataSuitabilityMA() to make MA plots and show that the data is unsuitable for constand
1053* 4cbf5dd adjust font sizes everywhere
1054* 5e7aa06 added Delta V boxplots to intraInterMAPlots()
1055* 3298ef0 refactoring the variable names in intraInterMAPlots() so that they remain distinguishable (different variable names for each case)
1056* 938f959 implemented boxPlot() for intraInterMAPlots
1057* af234e7 improvements for RHDPlot
1058* c34295f bugfix in compareDEAresults(): now uses the regular NON-log fold change
1059* edefcdb implemented RDHPlot() : relative difference histogram plot. measures |x-y|/max(x,y)
1060* a4e081f implemented compareDEAresults() for making MA plots of similar DEA analyses results (p-value and fold change)
1061* 99a2192 updated compareICmethods()
1062* 5a94d40 replaced MA plots mean by MAD (mean absolute deviation)
1063* 7b07202 Merge branch 'master' of ssh1.ulyssis.org:thesis
1064|\
1065| * 5af2f21 refactoring: split up intraInterMAPlots() into compareDIFFERENTconditions and compareIDENTICALconditions by means of 2 booleans
1066* | 44ab10b refactoring: split up intraInterMAPlots() into compareDIFFERENTconditions and compareIDENTICALconditions by means of 2 booleans
1067|/
1068* 806d53b added plot of last MA comparison to INTER and INTRA sections of intraInterMAPlots()
1069* f46addb implemented intraInterMAPlots
1070* eb3f500 refactoring: took the MA part out of MAPlot() and put into separate MA() function implementing intraInterMAPlots
1071* ee2ad53 refactoring: took the MA part out of MAPlot() and put into separate MA() function implemented
1072* 7f35d07 refactoring: getAllExperimentsIntensitiesPerCommonPeptide now returns a dataframe (with the intensity channels as headers) instead of a matrix
1073* e28e35a bugfix in getProteinDF when df doesnt contain 'Protein Descriptions': now automatically added (dummy)
1074* 96f9e0e started implementation of intraInterMAPlots() for making all intra and all inter-experimental MA plots for the Max data set
1075* 6fd7af5 bugfix in HCD: x and y labels were switched
1076* 52da0ee cherry-pick merge from abundances branch: compareAbundancesIntSN() implemented, getRTIsolationInfo now returns None if "representative first scan" or "RT" columns are missing.
1077* 7b7addd added hard code for COON_SN_nonormnoconstand
1078* a8832c9 added compareICmethods() for comparing the isotope corrections of proteome discoverer with constand++
1079* e831a4c TMT2ICM now returns the ICM with normalized rows.
1080* fe3fd13 collapsePTM is now enabled again!
1081* 52bea97 added separate COON_nonormnoconstand hard code in webFlow
1082* 7908e42 MAPlot title argument is optional
1083* 5315812 remove log files from repo
1084* 10cae0f getSortedDifferentialProteinsDF now sorts on adjusted p value instead of fold change, because p value already incorporates fold change.
1085* 9bccf90 applyDifferentialExpression now also remove proteins that have all nan values for a certain condition. Keep the removed ones in metadata
1086* 6f468ab applyDifferentialExpression now returns nan if p-value would be zero (one condition has all values missing)
1087* 9acc0e4 docu
1088* 13925af protein descriptions are identical across multiple peptides per protein: select only the first in the list (i.e. do not make a list of identical descriptions)
1089* 194a253 tons of stuff to make the MB_Bon dummy experiment working again
1090* 134e426 Merge branch 'master' of pow.ulyssis.org:thesis
1091|\
1092| * 6cb5e1b adding job folder in scripts dir
1093* | df26a01 generateReport() now also writes min and full sorted list of differentials to disk
1094|/
1095* 4b4451c disabled a test boolean
1096* 2f733d7 bugfix icm extention should not be in exportData call in transformICM()
1097* b38d7c6 pass channelNamesPerCondition to transformICM, not channelAliasesPerCondition
1098* 551c31b TMT2ICM docu: column order is changed
1099* 0675c84 bugfix ICMFile indentation in updateSchema()
1100* 909b86f bugfix ICMFile job path
1101* b92a286 bugfix if ICMFile would be None
1102* d29b71f TMT2ICM now reorders the columns and rows according to the nested list of channelAliasesPerCondition
1103* 1ad8183 implemented transformICM which takes a TMT ICM uploaded file and turns it into a real ICM and reorders the columns conform to the channelAliasesPerCondition by employing TMT2ICM
1104* 272a567 added COON_noISO exptype in webFlow()
1105* 4b5baf2 refactoring: moved TMT2ICM() and constructMasterConfigContents() from dataIO.py to webFlow.py
1106* 949afc5 removed empty entries from incompleteSchemaDict in parseSchemaFile also: this function should NOT be moved to web
1107* 8eb09aa TMT2ICM now returns floats in the ICM instead of percentages
1108* 1245849 bugfix in TMT2ICM: some observedChannels to not exist when Nplex is a subset of e.g. 10plex
1109* 23afb13 bugfix in getTMTIsotopicDistributions: index is now of type str
1110* 8e88de7 bugfix in getTMTIsotopicDistributions: modification must be inplace
1111* b9960bd docu change of TMT isotope table format
1112* 8e000fd implemented getTMTIsotopicDistributions to get a TMT table from file in the right dataframe format as require by TMT2ICM.
1113* 47d3fb5 implemented TMT2ICM for converting TMT isotope impurity distribution tables to isotopic correction matrices
1114* cbf3022 added doConstand option to abdundancesPCAHCD
1115* dfb470e implemented abdundancesPCAHCD function to analyze abundances of COON data instead of intensities or SNRs
1116* e9c9dda bugfix: MAPlot ignore nan and inf values when calculating var and mean
1117* 902e7a0 undo previous bogas bugfix; it was a stupid bracket
1118* 34c1cdb small bugfix in np import dtype
1119* 5162764 MA plot title (if provided) also gets mean M and var M
1120* c328775 added test code for making peptide-level MA plots
1121* 48743ed MAPlot() now takes a title as optinoal argument
1122* 0090db5 allExperimentsIntensitiesPerCommonPeptide is now added to reportFlow return and is included in analysisProcessingResults(Dump) so that it can be aesily passed to reportFlow
1123* e6cc84f added exportData for allExperimentsIntensitiesPerCommonPeptide so that we can do experiments or statistics on the peptide level
1124* ba9f9c5 fixed newlines at end of files
1125* 9191d6d fixed volcano plot code line location in reportFlow (in/outside if condition)
1126* 9ece1ac bugfix: baseJobConfig was removedfrom scripts directory, it is back now
1127* df15cfd stuff
1128* 26d1772 implemented functionality in compareIntensitySN to NOT do processing, but just constand only
1129* 278f076 bugfix import constand
1130* 13a674c changed compareIntensitySN pickle file location
1131* ec85d02 compareIntensitySN can now take two dataframes as argument
1132* 11646c1 refactoring: created processingFlow.py, analysisFlow.py, reportFlow.py and moved the functions processDf(), analyzeProcessingResults(), generateReport() from main.py to there. fixed imports
1133* f22c1b8 refactoring: main(): masterConfigFilePath->jobConfigFilePath
1134* 84db939 fixed devStuff compareIntensitiesSN() put test job files in job directory
1135* c267261 wrote some code (and immediately commented out) to check if allChannelAliases order remained consistent between getAllExperimentsIntensitiesPerCommonPeptide and getPCAPlot
1136* cdee3a2 TMT2ICM removed comments
1137* f3a4a73 getPCAPlot: added experiment label legend
1138* 7d44d1b hotfix: suddenly i get errors for DEA and volcano plot stuff if the number of conditions is not 2 ... should have happened before.
1139* 03fd20f comments
1140* ba688d1 getMarkers now returns a dict with markers per channel. Each experiment still has its own unique marker
1141* b2b8e2b getMarkers and getColours now get allChannelAliases as an argument instead of recalculating each time
1142* 3448d98 moved informative console prints in main into the doX booleans
1143* 1775996 figures should now be saved in larger size
1144* 8b2b524 figures should now be saved maximized
1145* 18a3b9c bugfix: visualizationsDict artifact
1146* b6a4abd makeHTML() has been adapted to "no more visualizationDict"
1147* 6951663 generateReport doesn't use a visualizationDict anymore. Just uses separate variable for each plot
1148* e1682cb exportData for type dataViz replaced by type fig. Doesn't accept dicts anymore
1149* e80eb29 hardcoded COON_SN_norm data
1150* 1a275d1 added informative console output strings in main
1151* 6b8b279 small bugfix when constand is not performed
1152* f80d4c0 hardcoded COON_norm experiment code
1153* e96f2fd implemented hardcoded switch in processDf to NOT use constand!
1154* ace15b6 added COON_SN to webFlow
1155* 77e4fe0 bugfix
1156* 0f4d0fc adapted getPCAPlot to use the new channelColorsDict
1157* 4e7bec8 getColours now returns a dict with 1 colour(=value) per channel(=key).
1158* dc1da5e fixed HCD colors
1159* 1c75882 increased PCA marker size
1160* 9936e26 apaprently some os.path's were still just path's????
1161* 0bbfbd4 implemented exportData for type viz: figures are saved both as png and as pickle so you can reopen and modify
1162* 3e5f9d9 implementing colors for getHCDendrogram changed distinguishableColours() cmap to jet
1163* bac7255 getColours docu more explicit
1164* d409c43 getHCDendrogram(): quality of life improvements + adding colours per condition
1165* 88d127a bugfix: removed Exception in getInput if the output path already exists, because I check this again in the main() method anyway.
1166* 96f2169 removed obsolete commented code
1167* faf46d5 moved getProcessingInput inside the same loop over the experiments as the doProcessing boolean. (the second loop was obsolete)
1168* f666813 refactoring: masterParams -> jobParams
1169* ac3f93e refactoring: specificParams -> processingParams
1170* 0b9ead4 implemented method to continue previous analysis: webFlow now takes argument previousjobdirName
1171* 623b6d2 bugfix in analyzeProcessingResults if noCorrectionIndices does not exist because isotopicCorrection_bool == False
1172* 4164b3c importDataFrame can now handle dtypes so that getWrapper -- CORRECTION: parseSchema() -- can specify "str" so as not to interpret 126 as a float 126.0, then to be cast to the wrong type of str value
1173* 5c470ca bugfix in updateConfigs and getProcessingInput: ICM can be None
1174* f892664 job name should be entered together with schema job path now also contains the job name
1175* bf31a29 webFlow now takes an exptype argument specifying which hardcoded experiments to use webFlow can now handle up to 4 hardcoded experiments added COON hard-code
1176* d9f2a8c webFlow can now handle up to 4 hardcoded experiments added COON hard-code
1177* 34baf43 made getProteinPeptidesDicts (and thus the whole workflow) independent of "# Protein Groups" column
1178* 90d5e72 implementing TMT2ICM (not finished yet)
1179* 80adc75 refactoring artifact
1180* 298fd35 implementing TMT2ICM converter (for web eventually) removed getIsotopicCorrectionsMatrix default path value
1181* b612f8f fixed resultsDumpFilename paths
1182* dea4320 refactoring: dataproc.py -> processing.py
1183* b935c9f refactoring: intensityColumnsPerCondition -> channelNamesPerCondition
1184* 733f21f removed intensityColumnsPerCondition from processingConfig. only needs intensityColumns
1185* 1d84b25 refactoring: masterConfig -> jobConfig; config -> processingConfig; baseConfig->baseProcessingConfig
1186* 52a3739 refactoring: getMasterInput -> getJobInput; getInput -> getProcessingInput
1187* 23e71ba changed all os.path.relpaths by os.path.abspaths in main()
1188* 8c6f2d2 bugfix: webFlow now corretly works on absolute paths but writes relative paths to config files
1189* 42662b5 bugfix: reverted accidental change of the masterConfig.ini in the scripts dir
1190* 500eda1 restructuring: path parameters in config.ini and masterConfig.ini should now be relative to the job path (specified by the parent path of the config file path) this has also been restructured in webFlow Also the hardcoded input paths in webflow are now gathered at the start of the file
1191* 5a5bf49 restructuring: path parameters in config.ini should now be relative to the job path (specified by the parent path of the config file path) this has also been restructured in webFlow Also the hardcoded input paths in webflow are now gathered at the start of the file
1192* 6094045 todos
1193* 4e6ff92 bugfix: some remaining filename_out replaced by jobname
1194* 319e1f0 bugfix: warnedYet was never set to true in isotopicCorrection
1195* 30788d8 bugfix: added logfile to generateReport call
1196* 9601af1 bugfix: something went wrong in the commit from 2 commits ago: mean of empty slice warning was still showing. also removed some warn() artifacts and replaced by logging.warning()s
1197* ad2fbca bugfix: maxRelativeReporterVariance artifacts
1198* 1d70696 found and silenced mean of empty slice spam source
1199* 720f86f removed maxRelativeReporterVariance and its flagging code which isnt being used
1200* 2fe6350 bugfix: masterparams does not require eName
1201* b79e847 moved the output dir generation back inside the experiment loop (analysis+result ones to respective other locations) but now correctly changes name on each loop --> each experiment processing has its own results subfolder
1202* 79f5f0a moved the output dir generation outside the experiment loop
1203* 7e99ad1 todo
1204* 5e2081a removed channelAliasesPerCondition from the specific params
1205* f2bcd63 bugfix: forgot to dumps channelAliasesPerCondition
1206* 79cb300 small bugfix: syntax
1207* a040881 bugfix: updateConfigs() wrongly wrote the original intensityColumnsPerCondition to the config.ini file instead of the aliases
1208* cf148fc bugfix: removed debugging artifact, replaced remaining warn()s by logging.warning()s
1209* b8d817c wrapper dataframe now imported as type str
1210* f9ed29f refactorign: dataIO.py: getDataFrame() -> getData()
1211* 4c3a95f bugfix path
1212* cc6fc3d Merge branch 'master' into webFlow (added logging)
1213|\
1214| * fdbab5f introduced logging replaces warnings.warn
1215* | a6b68e7 added channelAliasesPerCondition to getInput()
1216* | 95f9b07 added different subpaths for output (processing, analysis) and results refactoring: filename_out -> jobname
1217* | bcf3479 implemented parseDelimiter() to also allow for None delimiters
1218* | d99a310 delim_in is allowed to be None
1219* | b617f83 bugfix: updateConfigs now correctly discerns writing strings and non-strings updateMasterConfig too
1220* | fe48da8 added removedDataInOneFile_bool, header_in, collapse_maxRelativeReporterVariance to baseConfig
1221* | addf9ff updated dataIO.py: importDataFrame() so that it can let pandas automatically detect delimiters
1222* | 650d05f refactoring icm -> isotopicCorrection_matrix file_in -> data
1223* | d757577 fixed some wrong uses of write(), dumps() combinations
1224* | 32b9019 bugfix: forgot to write some newlines
1225* | 30a6a4b removed collapse_method from baseConfig
1226* | 3017c01 removed custom schema definition artifact from main()
1227* | bf33292 refactoring shadowing variables
1228* | 11cf39a bugfix: return masterConfigFile (is already full path)
1229* | 185b203 Merge branch 'master' into webFlow
1230|\ \
1231| |/
1232| * 92577cd updated importDataFrame: now drops lines that contain no values
1233* | 963a9c8 if no wrapper is uploaded, creates empty wrapper file to append to afterwards.
1234* | 3d0d044 implemented updateWrappers
1235* | 2da94b8 removed path_in from masterConfig and getInput
1236* | af65429 clarified webFlow steps
1237* | bb3f6d1 webFlow.py: made all paths full isntead of just filenames.
1238* | 17c970f bugfix indentation in updateConfigs(); better detection of [DEFAULT] line
1239* | f700337 separated webFlow() into a different file webFlow.py
1240* | b6fb9d9 added a newline first when appending to config files fout.write('\n') # so you dont accidentally append to the last line
1241* | 488ccc0 bugfix: config files should be appended with option a, not option w (write)
1242* | 4fa8bae newJobDir now returns an absolute path
1243* | c166d94 bugfix in date that is written to masterConfig
1244* | 2f1cb74 bugfix: dumps() does not write like config --> modify the write() operations so they write just like configparser
1245* | 2a7b319 bugfix in removal of files when uploadSchema fails; parseSchema argument path; updateConfigs;
1246* | 50a3113 bugfix in removal of files when uploadSchema fails and parseSchema argument path
1247* | 79f3ba1 Merge branch 'webFlow' of ssh1.ulyssis.org:thesis into webFlow
1248|\ \
1249| * | 1a9a538 implementing webFlow in main.py to simulate a web tool providing input added baseConfig.ini added ../jobs folder
1250* | | 93bfd4f implementing webFlow: implemented rest of step 3 and also step 4.
1251* | | e0ab61f implementing webFlow: put files in /jobs folder and implemented step 2 and part of 3 renamed baseConfig.ini
1252* | | f1150b7 implementing webFlow in main.py to simulate a web tool providing input added baseConfig.ini added ../jobs folder
1253| |/
1254|/|
1255* | 6dea623 bugfix: constructMasterConfigContents should now properly print tabs
1256|/
1257* 6b5a29e added masterInput parameter path_in
1258* 1bd3592 BUG NOT FIXED: updated constructMasterConfigContents which now STILL DOESNT handle delimiters correctly (visible format)
1259* 8261f64 updated constructMasterConfigContents which now handles delimiters correctly (visible format)
1260* a4691e0 import bugfix
1261* e8f1436 implemented constructMasterConfigContents
1262* a7bacc0 parseSchemaFile refactoring and docu
1263* 5722f29 implementing parseSchema
1264* 86543b2 multi-experiment bugfix in getAllExperimentsIntensitiesPerCommonPeptide
1265* a9b05cd stuff (implemented and commented out bad attempt at trying to make labels fit properly on the plot)
1266* beb31ae small bugfix
1267* 784700e hard bugfix in distinguishableMarkers. Fuck, be careful when assigning variables: use .copy(). I was destroying the matplotlib Markers archive.
1268* 2305b25 bugfixes in distinguishableMarkers: use the markers KEYS instead of VALUES
1269* cc8343e getPCAPlot is now multi-experimental and has the channelAliases as labels and each condition has its own colour
1270* 8a1c545 small bugfix
1271* f0c30ac bugfix in getColours maybe I should just use human readable for loops
1272* 3558cf2 small bugfix in getColours
1273* aa11362 refactoring: moved main.py:unnest to dataIO.py:unnest
1274* bd11de5 reimplemented getColours
1275* ce091c5 bugfix in getColours and getMarkers
1276* 4addd73 implemented getColours
1277* 9e0b0b8 implemented getMarkers
1278* bd65e3c implemented metadata['commonNanValues']
1279* 55bcb7e getAllExperimentsIntensitiesPerCommonPeptide() now also returns metadata['uncommonPeptides']
1280* 0b1c7fd refactoring in getAllExperimentsIntensitiesPerCommonPeptide(): pcadf -> peptidesDf
1281* 7948d87 refactoring: getAllExperimentsIntensitiesPerCommonPeptide
1282* 6753e1d fixed getPCA so that NaN values are set to zero 0.
1283* 079f755 refactoring fix
1284* d86ac28 implementing getAllExperimentIntensitiesPerPeptide
1285* 68678d6 implementing getPCAPlot for multiple experiments ... but first need to do analysis.py:getPCA for multiple experiments
1286* 2aef058 implemented distMarkers() for generating distinguishable markers
1287* a9d2467 added distColours to generate distinguishable colours
1288* 269949d getPCAPlot refactoring: now uses schema instead of intensityColumnsPerCondition
1289* 3ae2b9b stuff, bugfixes
1290* 4bb3570 added wrapper to schema
1291* 72ab8e0 removeMissing does not use "Quan Info" columns anymore to remove detections without quan values
1292* 801253c corrected removeObsoleteColumns docu.
1293* 7db38a8 ease of use
1294* 3d4716d implemented use of wrapper file wrapper.tsv moved input parameter modification getIsotopicCorrectionsMatrix and getWrapper away from the initial declaration of the variables (so that one first obtains them as a path string) commented out ICM checks and numpy import
1295* dcd9a20 refactoring: Identifying Node -> ~ Type implemented fix in fixFixableFormatMistakes if one wasnt using Identifying Node Type yet.
1296* ec541ab verified that flanking amino acids correction in fixFixableFormatMistakes works
1297* f5ee9e3 small bugfix: allChannelAliases definition
1298* 92c88a5 bugfix: cannot call .extend on [] expression. Just use "+" operator
1299* 793698d bugfix: conditionXIntensities should be Series
1300* 8adc989 small bugfix and undid some previous changes
1301* 6e63289 fixed getProteinDF() to properly handle multiple experiments. Now contains extra aggregation per experiment. now also interprets peptideIndices as MultiIndex
1302* 5a29483 fixed getProteinDF() to properly handle multiple experiments. Now contains extra aggregation per experiment.
1303* 91387ae bugfix in proteinDF definition in getProteinDF()
1304* dd8ddb5 bugfix in analyzeProcessingResult: definition of noCorrectionIndices
1305* ac74463 global variables are giving me shit. Commenting them all out. fixing the dependencies
1306* 800af64 global variables are giving me shit. Commenting them all out.
1307* 640201d fixed bug in getIntensities that was caused by a debugging artifact
1308* c2cb6f2 fixed bug in config2.ini intensityColumns
1309* 9b05a40 made 2 separate config files for separate experiments (on same data file)
1310* 5289504 bugfix in wrapper definition
1311* 38d7417 implemented applyWrapper() now works on df.columns instead of df
1312* 08e27fb bugfix in wrapper definition
1313* 0ad73e5 intensityColumnsPerCondition in specificParams now contains the values from channelAliasesPerCondition in masterParams['schema'] !!!
1314* d180792 bugfix in main(): global variables should be set in same loop as processDf calls
1315* f6b44cb small bugfix
1316* dc3a33a getPCA now omits NaN values
1317* a24f25f updated getPCA and getHC calls to use allChannelAliases as intensityColumns parameter in getIntensities call
1318* 3f79971 implemented unnest() in main.py
1319* 5d90f45 metadata is now multi-indexed
1320* fc0e71b fixed bug in getProteinDF call
1321* 61e5097 getIntensities now also accepts optional argument intensityColumns
1322* 72742be adapted getProteinDF to handle multi-experiment data
1323* af6f559 refactoring combineExperimentDFs(): no more fiddling around with intensityColumns or channelAliases: those names should be in-place from the beginning. BUT they are still called channelAliasesPerCondition and not intensityColumnsPerCondition in the masterParams !!!
1324* 9cf4a6d implementing multi-experiment support in analyzeProcessingResults. implemented combineExperimentDFs(): all dfs are merged together with multi-index into allExperimentsDF
1325* c8d29c3 implemented alias part of parseSchema
1326* 0d9713b implementing multi-experiment support in analyzeProcessingResults. all dfs are merged together with multi-index into allExperimentsDF
1327* 7994fdc possibly simplified code in getProteinDF
1328* f8b835d implementing analyzeProcessingResult for multiple experiments
1329* e04920d adapted analyzeProcessingResult to work with dict as processingResults input type
1330* bb36e9b modifying input and output paths
1331* 7ba2728 bugfixes
1332* f2ee668 small bugfixes
1333* c42343c small bugfix
1334* e0bc088 implementing experiment names via masterParams['schema']
1335* ca83a12 commented out parseSchema and refactoring getList -> parseExpression in getInput.py
1336* 29e42c8 schema in masterConfig is now a dict which is written automaticaly by the web interface, containing the column aliases, intensityColumnsPerCondition and the config file location.
1337* 6469814 implemented multiple experiments: * refactoring: parameter files_in -> file_in * splits params in specificParams and masterParams * added schema.tsv and schema masterConfig parameter * added parseSchema to dataIO.py * intensityColumnsPerConditionPerexperiment is now a dict of dicts * added masterConfig.ini which contains also the analysis+report master parameters * each experiment now has its own config file * each processing step is inside a big for loop, going over each experiment.
1338* 2e8a4ca added parameter identifyingNodes which now replaces masterPSMAlgo. Format: {"master": ["Mascot (A6)", "Ions Score"], "slaves": [["Sequest HT (A2)", "XCorr"]]}
1339* 90f60f2 fixed small bug and import
1340* 834cafa fixed imports
1341* 0b4c161 moved getInput to separate file. main() now takes an extra argument configFilePath
1342* 77a57e4 implemented fixFixableFormatMistakes
1343* 3e71715 implemented getDataFrame which now is superior to importDataFrame. made skeletons for fixFixableFormatMistakes and applyWrapper
1344* 4a4e2a8 Merge branch 'master' of ssh1.ulyssis.org:thesis
1345|\
1346| * ad9b266 added more code and MA plot for comparing intensity and S/N SNR
1347| * 2639794 commented adjustText import
1348* | e446751 makeHTML and HTMLtoPDF skeletons
1349* | 486d112 diffMinFullProteins contains the differences between the min and full protein lists
1350* | 622cdc8 refactoring: getSortedDifferentials -> getSortedDifferentialProteinsDF
1351* | 3eb7430 getSortedDifferentials now only returns important columns (protein, significant, description, fold change, p-value
1352* | 3987838 stuff
1353|/
1354* 4550714 fixed getSortedDifferentials
1355* 91eafc7 added todo
1356* 994bfb3 cleaned up some code in getVolcanoPlot
1357* 18f1d91 combined annotation code in getVolcanoPlot
1358* 5b837df two volcano plots are produced: one for minProteinDF, one for fullProteinDF
1359* 59ec516 added parameter labelVolcanoPlotAreas
1360* 05a9999 volcano plot now has labels, if enabled through list of booleans labelPlot
1361* 047fa48 PCA plot now has labels (channel)
1362* c56620d do not save visualization objects to disk (yet)
1363* 029b3d4 split getDataVisualization into getVolcanoPlot, getPCAPlot and getHCDendrogram
1364* f8f4fe3 giving dendrogram labels
1365* f96f9c8 intensityColumns parameter is now obsolete in the config file (replaced by intensityColumnsPerCondition)
1366* 5c2e559 getHC now first removes nan values
1367* 1ea4851 edited getHC because i misunderstood the instructions
1368* 3c6e7a1 bugfixes
1369* cd53954 added protein descriptions to proteinDF
1370* 8f52e2d implemented getSortedDifferentials which sorts the protein dataframe according to fold change
1371* 56b778f removed bloat from main()
1372* ab7ddb1 enabled dendrogram and added some todos
1373* 877acab applySignificance docu
1374* aedc568 enabled and fixed bug in getHC
1375* e73eefe implemented volcano plot
1376* ca6c821 added applySignificance() in analysis.py to create a column in the df which indicates the significance of the result
1377* 49d06bf implementing volcano plot
1378* 6944731 fixed PCA colouring
1379* 3a56041 refactoring df -> normalizedDf where applicable (like analyzeProcessingResults)
1380* a07be78 PCA plot now colours channels according to their corresponding condition PCA plot now has distinguishable colors for arbitrary number of conditions.
1381* 7a8d692 stuff
1382* e7e04e2 bugfix in analysisResults dump file save operation
1383* 826056d moved visualization save from analyseProcessingResults to generateReport
1384* 8896cff small bugfix
1385* 1ec2112 moved dataVisualization BACK to report.py because we don't need to save figures anyway, and this seems like a morelogical place. It is no problem to separate the PCA and the plotting of its results. Just passing the PCA results and stuff inside Python (=inside memory) is better than saving figures to disk and reloading them. Plus, you can now separate the visualization from the calculations, which is nice for debugging.
1386* a433ea9 bugfixes in dataVisualization
1387* a657722 added comment with instructions on saving figures
1388* 04635b9 added minimum PCA_components check (>=2)
1389* 2bdeb6b implemented PCA plot in dataVisualization
1390* e7860e6 getPCA was incomplete
1391* ef77ffe dataVisualization started implementing PCA plot
1392* a9f0880 refactoring: viz -> visualizationsDict
1393* 639c610 implementing dataVisualization hierarchical clustering dendrogram
1394* 3c6022a BUG in getHC: clustering causes segmentation fault SIGSEGV
1395* 170de3c getPCA now imputess missing values by 1/#channels for its calculation (it does not apply the imputation to the dataframe!)
1396* 7035072 bugfix in getPCA
1397* 9f73642 moved global parameter definition from processing() to main()
1398* 3343ec1 removed obsolete code
1399* e93deba implemented hierarchical clustering getHC()
1400* 4dbff67 added parameter PCA_components
1401* 931a534 unknown
1402* 29e6fa1 implemented PCA
1403* 3f9630a changed PCA implementation: now using scikit-learn package
1404* fd0321c implementing PCA
1405* c08a3a8 docu
1406* 37a988d moved visualizations back to the analysis part (because you have to save figures and stuff BEFORE you generate the report)
1407* 1aab95d bugfix: processing/analysis results should only be loaded if the booleans doAnalysis/deReport are actually True...
1408* 4f86148 added main.py:main() parameter doReport that can pick up on an earlier done analysis.
1409* 335ea46 created report.py which is to do the visualization and report generation
1410* fe2bf28 implemented test function compareIntensitySN() to get an idea of the difference between the use of intensities or S/N values
1411* cf83d14 small improvement + removed obsolete code
1412* a15d89a removed masked values from ttest output in applyDifferentialExpression
1413* d8acac6 small bugfix
1414* e206cda implemented applyFoldChange
1415* 7b50ff9 skeleton improvements
1416* 7798b3b small bugfix
1417* 14c4d31 dummy implementation of applyFoldChange
1418* da18c7c made results output nicer (changed column order, put protein back into the columns after the analysis is done)
1419* cda6cae in proteinDF the lists are now actually lists instead of series only the pvalues and adjusted pvalues are saved to file
1420* 76c81f4 bugfix: omit NA values in t-test
1421* 7135a92 small bugfixes
1422* 99b359d small bugfix
1423* b38903d bugfix in data exports (3rd return variable must be a dataframe)
1424* 69a3d5b small bugfix in data exports (no delim specified)
1425* d28586e small bugfix
1426* d5def89 bugfix in applyDifferentialExpression assignment
1427* cdbb831 small bugfix
1428* b6e75e0 bugfixes in appliDifferentialExpression bugfix in getProteinDF: now is a 3Nx1 array per condition instead of a Nx3 array
1429* 12eef3a bugfix: manual cast from int64index to list. very ugly
1430* 6571802 bugfixes in getProteinPeptidesDicts
1431* ae27d64 bugfixes in getProteinPeptidesDicts now warns when peptides without master protein accessions detected (numGroups == 0)
1432* a024f7a bugfix getProteinPeptidesDicts numGroups is of type int
1433* 8c71df8 refactogin proteinDF() -> getProteinDF()
1434* 22bcc06 removed illegal filename test case from files_path because it was annoying to debug.
1435* 4a430b5 import spelling mistake
1436* fb7bbfb abspath->relpath
1437* 46f7a21 exception clarification
1438* c5e2465 bug UNFIXED: this_proteinDF is empty when it enters applyDifferentialExpression
1439* ee7e490 bugfix in applyDifferentialExpression
1440* 8cbdac2 small bugfix in applyDifferentialExpression: .loc[] -> [] for new column definitions on dataframes
1441* 9f68421 bugfix in creation of empty dataframe
1442* 83f0b4e small bugfixes
1443* 5a20dad bugfixes (mostly apostrophes "" around integers in config file)
1444* e28eeba bugfix for multiple input files which is not completely implemented in main()
1445* d968e52 bugfix for multiple input files
1446* 7197772 config.ini update + bugfix
1447* a175032 bugfix
1448* 59b299a bugfix in json module import
1449* 75478e9 made data analysis callable without processing using a pickle python object written at end of (previous) data processing step
1450* 536d0bd fixed bugs: refactoring artifacts
1451* 93f1938 post-merge chaos
1452|\
1453| * c593240 added booleans doProcessing and doAnalysis to control which parts of the workflow to execute;
1454| * 00e311a implementing the workflow for allowing multiple experiments
1455* | f9c920f Merge remote-tracking branch 'origin/master'
1456|\ \
1457| * | 9c9f138 started elaborating analysis workflow
1458* | | 3db272e bugfix: didnt add parameters everywhere in dataIO
1459* | | bbc2f2f removed: dataIO.py:parseList() now json.loads handles lists, can also handle nested lists
1460* | | 289c763 added parameters alpha, FCThreshold, intensityColumnsPerCondition, pept2protCombinationMethod (mean/median)
1461* | | e519b4c bugfix: indexing in getProteinPeptidesDicts
1462* | | 6d182d2 fiddling around with imports
1463* | | ef7fba5 comments
1464* | | bf5e72a refactoring: foldChange -> applyFoldChange
1465* | | 75ba5fe refactoring: differentialExpression -> applyDifferentialExpression jsut modified the input df
1466* | | 992a3ec refactoring: shadowing variables
1467* | | f84f7bc removed obsolete code
1468* | | 33c337e refactoring: proteinIntensitiesPerConditionDF -> proteinDF implemented proteinDF
1469* | | ff38dd1 refactoring: setGlobals -> setProcessingGlobals
1470* | | ef352d4 implemented differentialExpression
1471* | | 3459319 started elaborating analysis workflow skeleton is as good as finished
1472|/ /
1473* | 7b2af4d setIntensities now handles both ndarrays and dicts
1474|/
1475* d788323 refactoring dataprep.py -> dataproc.py and very small bugfix
1476* 18ea9b8 refactoring: isotopicCorrectionsMatrix -> isotopicCorrection_matrix
1477* 2db00bf refactoring: remove_ExtraColumnsToSave -> removalColumnsToSave
1478* 7ac8d8f bugfix in setMasterProteinDescriptions: can now handle case where there is no Protein Descriptions colmun present (just does nothing but throw a warning in that case).
1479* c03b129 implemented analysis.py:getProteinPeptidesDicts() (including documentation)
1480* 005f81a refactoring: requiredColumns -> wantedColumns better reflects the nature of the list (see previous commit changes for context)
1481* 3f688b9 refactoring and re-implementation: getRequiredColumns -> removeObsoleteColumns because it now removes all except the wanted ones, instead of just selecting the wanted ones. This is because all the required columns may not be present but may still be non-essential (Xcorr, Ions score!) in other words they are not in noMissingValuesColumns
1482* 579e0a5 getBestIndices now still works if there is not PSM slave score column
1483* 9179fd7 Merge branch 'master' of ssh1.ulyssis.org:thesis
1484|\
1485| * e2ad4b8 Merge remote-tracking branch 'origin/master'
1486| |\
1487| * | dc45467 small bugfix
1488* | | b173477 changedconfig file tonot use Kurt's data
1489| |/
1490|/|
1491* | 4650f6b bugfix in isotopicCorrection: didnt properly return index of rows with NaN values
1492|/
1493* 64ab3f9 implemented warning and solutiong (pick first index in list of duplicates) for when there is no BestIndex for a list of duplicates.
1494* df8bb01 removed obsolete RT flagging code inside combineDetections()
1495* 0e2a99b added getNoIsotopicCorrection to analysis.py to record which detections received no isotopic correction.
1496* 5b70797 added duplicates column to getRTIsolationInfo
1497* 49b252d added metadata in main.py and getRTIsolationInfo in analysis.py
1498* ea08184 added date parameter
1499* e248ddb bugfix in combineDections: now returns a dict of intensities instead of just one vector with 6 values.
1500* 49637fc combineDetections for geometricMedian: rows should be normalized before applying geometric median, otehrwise the norm isnt conserved to one.
1501* c3a6496 bugfix in getIntensities for cases where df contained only one entry (then becomes a series)
1502* c016229 geometricMedian now throws away detections with missing values; 'mean' method now employs nanmean isntead of regular.
1503* 194c765 bugfix; removed intensityColumns dependency and used new version of getIntensities (see prev commit) instead
1504* f02c837 getIntensities can now handle a selection of indices
1505* b93d793 no collapsePTM allowed
1506* fe97df3 bugfix + clarification in setMasterProteinDescriptions
1507* 889971e implemented setMasterProteinDescriptions
1508* 9762a7b comments
1509* b630086 bugfix
1510* 2d43af0 bugfix
1511* efae067 refactoring: unusedDetectionsInOneFile -> removedDataInOneFile_bool
1512* 940ccbf bugfix in dataprep.py:setGlobals
1513* 16a8d69 added functionality to save removedData as one file
1514* b4b1a6e simplified code redundancy in undoublePSMAlgo and added unusedDectionsInOneFile_bool parameter
1515* 2e5f8db added noMissingValuesColumns and remove_Extracolumns parameters so that very few columnsToSave entries are hardcoded (only those who will never change)
1516* fc3c63f removed collapseRT_bool
1517* 3a146fd fixed bug in getDuplicates for collapse PTM: annotated sequence is now first .str.upper()-ed before groupBy annotated sequence.
1518* c5c7d95 fixed off-by-one error in getRepresentativesDf --> BIG BUGFIX off-by-one + race condition during data manipulation = results non-deterministic
1519* 03fbe13 fixed off-by-one error in getRepresentativesDf --> BIG BUGFIX off-by-one + race condition during data manipulation = results non-deterministic
1520* e3738e5 rewrote getIntenseIndices stuff to also use a dict
1521* 58fca2c rewrote getIntenseIndices --> getIntenseIndicesDict (self-explanatory) { bestIndices : intenseIndices }
1522* 0e38acc some refactoring even more df[ --> df.loc[ and refactoring to avoid shadowing and rewriting collapse() so that bestIndices --> bestIndicesDict is properly used
1523* 69f15f1 refactoring even more df[ --> df.loc[
1524* 9ea39ae refactored getBestIndices (code redundancy)
1525* 73724c6 refactoring to avoid variable name shadowing
1526* bd02c2e seems to be a VERY WEIRD bugfix ( you know, the bug where the output of the program was randomized and suddenly the dataframe had itself inside of itself on some index and on the other indices just 22 times "somestring" while it looked completely normal on inspection through the debugger but not when you called its indices... ) where I now changed representativesDf['Degeneracy'] ------> representativesDf.loc['Degeneracy'] and it seems to be solved
1527* e26fdbc small possible bugfix in getRepresentativesDf
1528* ee874c1 assert statement
1529* 6de2989 simplified removedData construction (direct from df instead of list of dataframes per duplicates group)
1530* 22d8695 documented that the order of adding representatives in getrepresentativesDf is important.
1531* 7939bc3 bugfix in groupByIdenticalProperties: now also checks if the candidate duplicates group consists ofmore than 1 member AFTER all groupby's have been done. Also, now uses filter() ahead of the for-loop instead of an if statement inside the for-loop.
1532* 5ef40aa removed obsolete code that replaces groupByIdenticalProperties
1533* 70b51dc small performance increase in groupByIdenticalProperties
1534* f84078d badConfidence->confidence (removedData name)
1535* 0ca1e09 fixed performance issue involving chained assignment
1536* c8c2729 isotopicCorrections now doesnt try to correct detections that have NaN values in their intensities. Gives a warning if this occurs at least once
1537* 2ceddbc removed profiler inside code, using profiler.py
1538* 52d12ee added profiler
1539* ed7af8b bugfixes in undoublePSMAlgo
1540* 726f925 bugfixes in undoublePSMAlgo
1541* 7d6889a removedData in collapse: now saves representative first scan instead of index (because index is unknown to the user)
1542* b0a2c5f removed some SettingWithCopyWarning
1543* fd23548 exportData can now handle dataframes in removedData object and write them separately to dsik.
1544* ff1649d bugfix in: unified all removedData formats. collapse function's removedData is now one dataframe containing all collapsed detections along with their parent ID.
1545* 3fd52dd small bug in undoublePSMAlgo assertion
1546* e3e5b73 unified all removedData formats. collapse function's removedData is now one dataframe containing allcollapsed detections along with their parent ID.
1547* 41ffb86 removedData in undoublePSMAlgo is no longer a tuple (master PSMAlgo clear from removedData PSM score column name
1548* 2d1b8ab removedData in undoublePSMAlgo is no longer a tuple (master PSMAlgo clear from removedData PSM score column name
1549* b7a797f bugfix in sanity check
1550* 45cdf0e bugfix in getRepresentativesDf
1551* 4124ba2 bugfix in getRepresentativesDf
1552* d186ba9 main() now takes testing and writeToDisk as argument booleans. implemented testDataComplementarity
1553* 6e9968c documentation
1554* 6cbb2f7 bugfix in exportData
1555* ddb32ed exportData documentation
1556* 60633bb config filename extension removal
1557* 734080d elaborated exportData for objects (like removedData)
1558* 185c6bd bugfixes; getBestIndices
1559* aa968c3 bugfixes; getrepresentativesDf
1560* 47a0134 bugfixes; groupByIdenticalProperties
1561* d84e1ac bugfixes; made toDelete in undobulePSMAlgo into Set
1562* 8a45cbf documented getRepresentativesDf fixed indentation bug and renamed indices->"group of indices" in duplicateLists explanation
1563* b5740c8 documented getIntenseIndices
1564* b33a58f updated getDuplicates:groupBy... documentation and documented getBestIndices
1565* 28b2a8a implemented degeneracy propagation in getRepresentativesDf()
1566* 562d48c refactoring: scanDict -> byFirstScanDict
1567* c173385 elabroated combineDetections(), removed getNewIntensities()
1568* 8c3bba6 added locations for flagging and implemented flagMaxReporterVariance in combineDetection() in collapse() in collapse.py
1569* 0f359b4 cleaned up some old code in getRepresentativesDf
1570* 77ad79d implemented the saving of the removedData in collapse.py:collapse()
1571* c679b9f implemented getRepresentativesDf()
1572* 505c7af BIG CHANGES: rewriting collapse() and all its functions. representatives are now a COPY of the best matching of the duplicates, and their intensities are modified if method is not bestMatch.
1573* 76833da implemented getDuplicates() to use the groupby function. Also made it compacter using grouByIdenticalProperties(), but also left old code intact and implemented switch youreFeelingLucky boolean
1574* d5d295c reworked undoublePSMAlgo to use the groupby function
1575* 9f34665 added masterPSMAlgo to collapse.py:getRepresentative()
1576* fce5573 removeBadConfidence documentation
1577* 236a228 added degeneracy column when collapsing
1578* 40dfdf0 refactoring: columnsToSave -> collapseColumnsToSave for collapse.py usage everywhere
1579* dcfc252 fixing some bugs in possible collapse function combinations
1580* 657f818 added and implemented removeMissing in dataprep.py
1581* 69cd246 added and implemented removeMissing in dataprep.py
1582* 551059e refactoring: made columnsToSave a parameter in config.ini
1583* b22627c unified colsToSave for all collapse cases
1584* 087db31 fixed bug in configparser for # symbols holy shit why do people put # signs in their legit output
1585* 59e5ded added setIntensityColumns in dataprep.py elaborated selectEssentialColumns in dataprep.py included essentialColumns and intensityColumns as parameters in config.ini added parseList in dataIO.py
1586* 9985f0f added removeBadConfidence function along with its two parameters
1587* 5bed16f fixed some import bugs due to refactoring
1588* b1d25e3 refactoring: moved all collapse-related functions to collapse.py from dataprep.py and modified their documentation added selectEssentials()
1589* e513cfe refactoring: collapsePSMAlgo -> undoblePSMAlgo
1590* e906530 fixed bug in collapse function calls in main()
1591* bfcd94f refactoring: collapsePSMAlgo_master -> masterPSMAlgo
1592* eb0bca9 refactoring: generalized all collapse functions into one collapse function (except collapsePSMAlgo because its not a real collapse) added function getRepresentative, made combineDetections more explicit
1593* c844844 refactoring: generalized all collapse types/methods and centermeasures into one collapse_method parameter, used by all collapse functions.
1594* 83c63c0 refactoring: created getNewIntensities() function which replaces the identically named functions in the body of the collapseXXX() functions.
1595* be7aec3 refactoring: created collapse() function which now handles almost all of the commands (NOT function definitions) in the collapseXXX() functions.
1596* 21c8b65 implementing new dataprep.py:collapsePSM()
1597* 738a4fd implemented sanity check: main() checks whether there are still identical (!= duplicate) peptides after collapsePSMAlgo()
1598* 90d2f8d implemented sanity check: main() checks whether there are still duplicate sequences after all collapses have been applied
1599* 5bf06da implemented sanity check: checkTrueDuplicates of collapseCharge() now asserts that you ran collapseRT() implemented sanity check (test): also assumes you ran collapsePSMAlgo()
1600* 8567349 decided on collapseRT method and centerMeasure structure. Implemented barebones version in collapseRT() and created new fucntion combineDetections() which contains the centerMeasure part.
1601* 2531478 decided on collapseRT method and centerMeasure structure. Implemented barebones version in collapseRT().
1602* 5ebfb81 refactoring: collapseRT_centerMeasure_addition -> collapseRT_centerMeasure
1603* 0a36365 created MS2IntensityDoesntMatter() dev method
1604* 40c49ac refactor collapseRT_centerMeasure_intensities -> collapseRT_method
1605* d903332 refactor collapseRT_centerMeasure_reporters -> collapseRT_centerMeasure_addition
1606* bc92006 added missing removedData assignments in main()
1607* 429a0ae added collapseCharge_bool dependency on collapseRT_bool in getInput()
1608* 56c3775 refactor: .iloc->.loc for fear that positional instead of label-based indexing might fck things up, although it looks like it doesn't.
1609* 1de8a37 tested isotopicCorrections(): it works
1610* 3d167a9 bugfix: isotope impurities should be imported as float64 instead of int64
1611* bdcadce refactoring: channels->reporters
1612* b5e9e84 bugfix in getIsotopicCorrectionsMatrix()
1613* c0bf5c2 bugfix in getNewIntensities()
1614* b131e84 isotopic corrections documentation
1615* 99eab74 stuff
1616* 482ab88 implemented dataprep.py:isotopicCorrections()
1617* 65ab8ce added ICM check
1618* 96526e6 moved isotopiccorrectionsmatrix to a separate file.
1619* 2c978bf bugfix in removedData
1620* 1c4a468 implementation collapseRT() and refactoring maxRelativeReporterVariance
1621* a067f0b moved getDuplicates() OUTSIDE the collapseCharge() function so that it can also be used by the collapseRT() function. It now also needs the dataframe and a checkTrueDuplicates() function.
1622* 48a0c33 modified the way collapseCharge retains deleted data: associated each duplicate row with its matching firstOccurrence using a dict
1623* 9710f37 documentation minor name change
1624* a8fbefc name change updateFirstOccurrences() -> getNewIntensities() documentation
1625* c70986d modified updateFirstOccurrences() and collapseCharge() to produce a dict for setIntensities() input
1626* c674ddf implemented setIntensities using a dict as input
1627* 4fb7b85 implementing collapseCharge() but this is a freaking b*tch there are still issues with a) how to detect peaks that ought to be collapsed and b) how to combine the intensities of the duplicates
1628* 4ed60b8 implementing collapseCharge() but this is a freaking b*tch
1629* aa01df2 added ext2delim and delim2ext functions. streamlined importDataFrame() using those functions
1630* 0b06dce added ext2delim and delim2ext functions. streamlined importDataFrame() using those functions
1631* 99114f8 fixed bugs in dataprep.py:collapsePSMAlgo() changed dataIO.py to use the config.get-methods instead of castin, because the getboolean() parser is very neat and bool() casting is dumb.
1632* 074091b fixed bugs in dataIO.py:getInput()
1633* 49ce67e fixed bugs in dataIO.py:getInput() but still need to fix many (apply variable changes to params dict)
1634* 47cd3c0 dataIO.py:getInput() now uses a config parser created config file config.ini
1635* f4b3bc0 Merge branch 'master' of ssh1.ulyssis.org:thesis
1636|\
1637| * 20b3daa implemented collapsePSMAlgo() (still buggy)
1638* | 220d2f6 fixed a refactoring error
1639* | 996568d removed some whitespace
1640* | 95417cb tweaked performanceTest to not use dataIO.py
1641* | de705de added isotope impurities matlab scripts
1642* | d150615 cleaned up as much testing stuff from main method as possible
1643* | 5cb48ec tested for isotope impurities correction by constand or not (NO)
1644* | 33466e4 added dataType param to exportData
1645* | 366f97f implemented collapsePSMAlgo() (still buggy)
1646|/
1647* c35a65c added test_dataprep.py (barebones)
1648* 9cb79b5 added test_dataprep.py (barebones)
1649* 5c788dd isotopicCorrection: do not output corrected intensities to user.
1650* 4519ac7 implemented removeIsolationInterference() and removed addColumns in favor of a dict removedData.
1651* 0cce977 stuff
1652* 6f948f2 skeleton of peptide to protein map in analysis.py:mapPeptToProt()
1653* ef4d456 a LOT of bugfixes in the skeleton / parameter chaos
1654* 67059ca forgot isolationinterference parameters in dataIO.py
1655* 9977555 small bugfix
1656* 08ee752 for now I think all the necessary parameters are included
1657* 420c7bf added DEFoldThreshold
1658* b534abf added collapseRT_maxRelativeChannelVariance parameter (damn these names are getting long.....)
1659* 87bcffd addid removeIsolationInterference (barebones)
1660* 974bf0e modified exportData to handle both a path plus a filename instead of only a full filepath
1661* 32112c6 removed the importData() function (importDataFrame() does all that is needed).
1662* a1b1ffd further explicited workflow skeleton and function skeletons (added parameters). added function (barebones) to add columns according to workflow format for retaining to-be-deleted info. also added more parameters checks as parameters are added.
1663* 3e7a24e renamed selectIntensities->getIntesities moved getIntensities to dataprep (from dataIO) added setIntensities (barebones)
1664* fed4b2c stuff
1665* d37ca56 dataIO.py:getInput() now returns the parameters inside a dict "params"
1666* eecfc9a built the entire workflow skeleton inside the main.py:main() and made barebones files (analysis.py, dataprep.py) with the required functions.
1667* 62854ae added barebones test file test_dataIO.py
1668* b9c330a fixed basic unit tests for test_constand
1669* a2146b6 added header_in argument to importData and importDataFrame in dataIO.py
1670* 25091e4 made basic test_constand
1671* 66848f8 removed nan values in returned convergenceTrail
1672* 1a555f6 added checks to dataIO.py and small refactoring
1673* a04154a added checks to dataIO.py and small refactoring
1674* 86ddd4e stuff
1675* e23e319 docu of constand.py completed (for now)
1676* 9352038 docu of dataIO.py completed (for now)
1677* a4be0c2 started documenting (dataIO.py)
1678* a35cee4 initialized test file for constand.py:constand() created dataIO.py and migrated code from main.py into it
1679* ced3d37 added exportData function
1680* 3dda0ab added exportData function
1681* 702f91d mail.py:selectIntesities is working now
1682* 83f220a fixed bug in importDataFrame elaborated importDataFrame (added more security checks) elaborated getIntensities (not finished yet)
1683* f519a5e added main.py:selectIntensities (empty for now) which should extract the intensities from the dataFrame
1684* b2cf3e9 elaborated profiler function (though some more tweaking needed + need to add line_profiler switch, now only has cProfiler)
1685* de315fd elaborated importData function to contain test stuff (does not call importDataFrame anymore)
1686* 5978b24 added profiler.py for testing the performance
1687* b6c4d8f elaborated main.py:importData and renamed to importDataFrame. importData should become a different function.
1688* 19cbd1d (moved from constand_matrix_mult) aesthetics + main.py:importDataFrame->importData
1689* 2d41a4d aesthetics
1690* 68e570a added performanceTest method so that i dont have to copy paste that code each time.
1691* 5ebb027 bug in constand.py:convergencetrail gefixt; NaN waarden en overeenstemming met MATLAB output succesvol getest.
1692* d975ae1 import_MB....m script elaborated (is eigenlijk gebeurd maar is toen niet bij in de commit geraakt precies)
1693* 014d767 bugs in constand.py gefixt. Algoritme werkt nu met Ri en Si per iteratie ipv enkel het totale product R en S. (maar R en S worden wel gaandeweg berekend zodat het niet nodig is extra memory te voorzien)
1694* 92c9ede implemented constand.py but it still has bugs; elaborated the main skeleton (getInput etc)
1695* 7ffbdb4 more matlab test scripts, created constand.py
1696* aa18911 main file skeleton
1697* 9561e0d initial commit