Global ETD Search

31	Exploiting whole-PDB analysis in novel bioinformatics applications Ramraj, Varun January 2014 (has links) The Protein Data Bank (PDB) is the definitive electronic repository for experimentally-derived protein structures, composed mainly of those determined by X-ray crystallography. Approximately 200 new structures are added weekly to the PDB, and at the time of writing, it contains approximately 97,000 structures. This represents an expanding wealth of high-quality information but there seem to be few bioinformatics tools that consider and analyse these data as an ensemble. This thesis explores the development of three efficient, fast algorithms and software implementations to study protein structure using the entire PDB. The first project is a crystal-form matching tool that takes a unit cell and quickly (< 1 second) retrieves the most related matches from the PDB. The unit cell matches are combined with sequence alignments using a novel Family Clustering Algorithm to display the results in a user-friendly way. The software tool, Nearest-cell, has been incorporated into the X-ray data collection pipeline at the Diamond Light Source, and is also available as a public web service. The bulk of the thesis is devoted to the study and prediction of protein disorder. Initially, trying to update and extend an existing predictor, RONN, the limitations of the method were exposed and a novel predictor (called MoreRONN) was developed that incorporates a novel sequence-based clustering approach to disorder data inferred from the PDB and DisProt. MoreRONN is now clearly the best-in-class disorder predictor and will soon be offered as a public web service. The third project explores the development of a clustering algorithm for protein structural fragments that can work on the scale of the whole PDB. While protein structures have long been clustered into loose families, there has to date been no comprehensive analytical clustering of short (~6 residue) fragments. A novel fragment clustering tool was built that is now leading to a public database of fragment families and representative structural fragments that should prove extremely helpful for both basic understanding and experimentation. Together, these three projects exemplify how cutting-edge computational approaches applied to extensive protein structure libraries can provide user-friendly tools that address critical everyday issues for structural biologists. 572.80285
32	Genetics of ankylosing spondylitis Karaderi, Tugce January 2012 (has links) Ankylosing spondylitis (AS) is a common inflammatory arthritis of the spine and other affected joints, which is highly heritable, being strongly influenced by the HLA-B27 status, as well as hundreds of mostly unknown genetic variants of smaller effect. The aim of my research was to confirm some of the previously observed genetic associations and to identify new associations, many of which are in biological pathways relevant to AS pathogenesis, most notably the IL-23/T<sub>H</sub>17 axis (IL23R) and antigen presentation (ERAP1 and ERAP2). Studies presented in this thesis include replication and refinement of several potential associations initially identified by earlier GWAS (WTCCC-TASC, 2007 and TASC, 2010). I conducted an extended study of IL23R association with AS and undertook a meta-analysis, confirming the association between AS and IL23R (non-synonymous SNP rs11209026, p=1.5 x 10-9, OR=0.61). An extensive re-sequencing and fine mapping project, including a meta-analysis, to replicate and refine the association of TNFRSF1A with AS was also undertaken; a novel variant in intron 6 was identified and a weak association with a low frequency variant, rs4149584 (p=0.01, OR=1.58), was detected. Somewhat stronger associations were seen with rs4149577 (p=0.002, OR=0.91) and rs4149578 (p=0.015, OR=1.14) in the meta-analysis. Associations at several additional loci had been identified by a more recent GWAS (WTCCC2-TASC, 2011). I used in silico techniques, including imputation using a denser panel of variants from the 1000 Genomes Project, conditional analysis and rare/low frequency variant analysis, to refine these associations. Imputation analysis (1782 cases/5167 controls) revealed novel associations with ERAP2 (rs4869313, p=7.3 x 10-8, OR=0.79) and several additional candidate loci including IL6R, UBE2L3 and 2p16.3. Ten SNPs were then directly typed in an independent sample (1804 cases/1848 controls) to replicate selected associations and to determine the imputation accuracy. I established that imputation using the 1000 Genomes Project pilot data was largely reliable, specifically for common variants (genotype concordence~97%). However, more accurate imputation of low frequency variants may require larger reference populations, like the most recent 1000 Genomes reference panels. The results of my research provide a better understanding of the complex genetics of AS, and help identify future targets for genetic and functional studies. 616.73

Search results

Exploiting whole-PDB analysis in novel bioinformatics applications

Genetics of ankylosing spondylitis