Publications
Projects
login
About
login:
password:
Forgot your password?
Investigation of economical and practical aspects of commer...
Investigation of economical and practical aspects of commercial cloud computing for Life Sciences
Publication type:
mastersthesis
Zusammenfassung:
In Life Sciences werden viele Rechenressourcen für das Prozessieren und Analysieren der erzeugten Datenmengen benötigt, beispielsweise für die Spektrum-Sequenz Zuordnungen in Shotgun-Proteomics. Diese Rechenressourcen sind oftmals massiv, werden aber meist nur kurzzeitig benötigt. Blade Systeme oder Rechen-Cluster bestehend aus mehreren hundert oder auch tausenden Rechnern, verwaltet mit Job Scheduling Sytemen, wie Portable Batch System oder Sun Grid Engine sind oftmals Lösungen, die man an akademischen Einrichtungen vorfindet. Mit den immer weiter wachsenden Datenmengen verschiedener Messinstrumente stossen diese lokalen Infrastrukturen jedoch an Kapazitätsgrenzen, da ihre Kapazität nicht "on-demand" ausgebaut werden kann. Daher stellen kommerzielle globale Grossrechenanlagen, wie das Amazon EC2 System, eine interessante Alternative für zukünftige Rechenengpässe dar. Im Rahmen dieser Diplomarbeit soll untersucht werden, inwieweit sich Cloud-Computing Lösungen für die parallele Prozessierung realer Datensätze aus der Bioinformatik eignen. Speziell, sollten hier wirtschaftliche sowie technische Aspekte bezüglich der Portierung von Algorithmen aus den Bereichen "Shotgun-Proteomics" sowie "Deep-Sequencing" auf Cloud-Computing untersucht werden. Dazu soll ein Konzept erarbeitet werden, welches dem Benutzer erlaubt, schnell Aussagen darüber zu treffen, welche Art von Rechensystem in Abhängigkeit der Anwendung und der Grösse der Eingabedaten, wirtschaftlich und sinnvoll genutzt werden kann. Insbesondere soll die Diplomarbeit Aufschluss geben über Verfügbarkeit, Preis-Leistungsverhältnis (Rechenkosten, Input/Output Datentransfer- Kosten), Skalierbarkeit und Portabilität zu exisitierenden Lösungen (insbesondere UZH Schroedinger Cluster / FGCZ Cluster). Um das Ziel zu realisieren ist es notwendig, dass der im Rahmen dieser Arbeit entwickelte Software-Prototyp mit unserem Daten Management Systems kommuniziert.
Authors:
Aleksandar Markovic
Abstract:
Many processes in life sciences involve the handling and processing of large data sets and require hence large amounts of computing power. Peptide-spectrum assignments in Shotgun-Proteomics constitute such an example. A lot of academic institutions install thus large cluster systems, consisting of hundreds or even thousands of computers administered by Sun Grid Engine (or similar approaches) for solving such problems. However, these computing resources, that are usually very massive, are often used for a short period of time. As the data amounts of various measurement instruments increase, those local infrastructures are pushed to their limits because the compute capacity can not be extended "on demand". In this context, commercial cloud computing solutions like Amazon EC2 could provide an interesting, alternative solution for future compute power shortages. In this thesis, we investigate the appropriateness of such cloud computing solutions for parallel processing of real datasets in bioinformatics. In particular, the economical and technical aspects of porting the algorithms from the domain of "Shotgun-Proteomics" and "Deep-Sequencing" to cloud computing will be investigated. Furthermore, a concept that would allow the user to quickly find an answer to the question which computing system is more suitable in a particular situation depending on the application and the amount of input data will be developed. Consequently, this work will give information about availability, value for money (computing costs, input/output data transfer costs), scalability and portability both for existing solutions (UZH Matterhorn Cluster / FGCZ Cluster)) and for the cloud computing approach. Based on this data one can find out at what point commercial cloud computing is cheaper than "local cluster computing" by considering purchase costs, maintenance, staff costs, cooling costs and so on. To achieve this purpose, a lot of simulations on the cloud computing side need to be performed for jobs that are already processed in the local cluster. These simulations would allow to collect information about correctness, running time, job size and other interesting features that would enable to discern the advantages of one method or the other. The software prototype, developed in the context of this thesis will be able to communicate with the FGCZ data management system, where the data sets as well as their processing results reside.
Title:
Investigation of economical and practical aspects of commercial cloud computing for Life Sciences
Year:
2010
month:
March
school:
University of Zurich
group:
csg
actions