Global ETD Search

1	Subfamily classification of the Defensin gene superfamily Shikhagaie, Medya January 2004 (has links) Defensins are small cysteine-rich, cationic peptides that play an essential role in the innate immune system of virtually all life forms, from insects and plants to amphibians and mammals. Defensins are mainly an innate immunity element, exhibiting antibacterial activities by disrupting the cell membrane of a wide range of organisms (Cole et al. 2002). Defensins also affect certain adaptive immune responses, including enhancing phagocytosis, promoting neutrophil recruitment, and enhancing the production of proinflammatory cytokines. The aim of this thesis is to make a comprehensive and accurate subfamily classification of the defensin gene family, primarily by using a library of Hidden Markov Models (HMMs). In this project the subfamily classification of the defensin gene family is primarily based on a constructed library of HMMs. Results: Sets of known defensins were organized in placed in 84 clusters using the clustering and alignment tool, FlowerPower. The clusters were further classified as mammalian alpha- or beta-defensins, plant defensin, insect defensin and defensin MGD. This classification was based on significant cluster hits against the Structural Classification of Proteins (SCOP) database and species distribution. Based on the relative positions of disulfide bonds and constructed Multiple Sequence Alignments (MSAs) some sequences were classified as belonging to the sperm– and theta-defensin subfamilies. Compared to PFAM’s classification of defensins, the subfamily classification presented here is more informative. The library of HMMs has been made public via a web server that was used to automatically score and analyze input sequences against the created database of HMMs. This database and web server are expected to be useful to researchers working on various aspects of defensin action. Defensins FlowerPower SCOP HMMs Bioinformatics Bioinformatik
2	Subfamily classification of the Defensin gene superfamily Shikhagaie, Medya January 2004 (has links) <p>Defensins are small cysteine-rich, cationic peptides that play an essential role in the innate immune system of virtually all life forms, from insects and plants to amphibians and mammals. Defensins are mainly an innate immunity element, exhibiting antibacterial activities by disrupting the cell membrane of a wide range of organisms (Cole et al. 2002). Defensins also affect certain adaptive immune responses, including enhancing phagocytosis, promoting neutrophil recruitment, and enhancing the production of proinflammatory cytokines.</p><p>The aim of this thesis is to make a comprehensive and accurate subfamily classification of the defensin gene family, primarily by using a library of Hidden Markov Models (HMMs). In this project the subfamily classification of the defensin gene family is primarily based on a constructed library of HMMs. Results: Sets of known defensins were organized in placed in 84 clusters using the clustering and alignment tool, FlowerPower. The clusters were further classified as mammalian alpha- or beta-defensins, plant defensin, insect defensin and defensin MGD. This classification was based on significant cluster hits against the Structural Classification of Proteins (SCOP) database and species distribution. Based on the relative positions of disulfide bonds and constructed Multiple Sequence Alignments (MSAs) some sequences were classified as belonging to the sperm– and theta-defensin subfamilies. Compared to PFAM’s classification of defensins, the subfamily classification presented here is more informative. The library of HMMs has been made public via a web server that was used to automatically score and analyze input sequences against the created database of HMMs. This database and web server are expected to be useful to researchers working on various aspects of defensin action.</p> Defensins FlowerPower SCOP HMMs Bioinformatics Bioinformatik
3	Reconhecimento e predição de promotores procarióticos: investigação de uma metodologia in silico baseada em HMMs Reis, Adriana Neves dos 03 March 2005 (has links) Made available in DSpace on 2015-03-05T13:53:45Z (GMT). No. of bitstreams: 0 Previous issue date: 3 / Universidade do Vale do Rio dos Sinos / A expressão dos genes em procariotos é desencadeada quando a enzima RNApolimerase interage com uma região adjacente ao gene, chamada de promotor, onde se encontram os principais elementos regulatórios do processo de transcrição. Apesar do crescente avanço das técnicas experimentais em biologia molecular, caracterizar e identificar um número significante de promotores, presentes em um dado genoma, continua sendo uma tarefa demorada e cara. Abordagens in silico são bastante utilizadas para reconhecer essas regiões em procariotos. Entretanto, além do alto número de falsos positivos obtidos, elas enfrentam a inexistência de um número adequado de promotores conhecidos para identificar padrões conservados entre as espécies. Logo, um método criterioso e confiável para predizêlos em qualquer organismo procariótico ainda é um desafio. Esta dissertação propõe um protocolo de uso de hidden Markov models (HMMs) que emprega Estimação de Limiar de Decisão (ELD) e Análise de Discriminação (AD) neste problema. Quatro espécie / Gene expression on prokaryotes initiates when the RNA-polymerase enzyme interacts with DNA regions called promoters. In these regions are located the main regulatory elements of the transcription process. Despite the improvement of in vitro techniques for molecular biology analysis, characterizing and identifying a great number of promoters on a genome is a complex task. In silico approaches are usually employed to recognize theses regions on prokaryotes. Nevertheless, the main drawback is the absence of a large set of promoters to identify conserved patterns among the species. Hence, a in silico method to predict them on any species is a challenge. This work proposes a protocol to use hidden Markov models (HMMs) methodology with Decision Threshold Estimation and Discrimination Analysis on this problem. Four prokaryotic species are investigated (Escherichia coli, Bacillus subtilis, Helicobacter pylori e Helicobacter hepaticus). The influence of different aspects in the recognition and prediction are examined: Ciências Exatas e da Terra procariotos reconhecimento de padrões HMMs promotores HMMs pattern recognition prokaryotes promoters
4	Uma abordagem integrada para a construção e utilização de HMMs de perfil para análises genômicas e metagenômicas / An integrated approach for the construction and application of profile HMMs for genomic and metagenomic analyses. Kashiwabara, Liliane Santana Oliveira 02 August 2019 (has links) HMMs de perfil são um método poderoso para modelar a diversidade de sequências biológicas e constituem uma abordagem muito sensível para a detecção de ortólogos remotos. Uma potencial aplicação de tais modelos é a detecção de vírus emergentes e novos elementos genéticos móveis. Nosso grupo desenvolveu recentemente o GenSeed-HMM, um programa que emprega HMMs de perfil como sementes para montagem progressiva de genes-alvo, utilizando tanto dados genômicos como metagenômicos. No presente trabalho foi desenvolvido o TABAJARA, um programa para o desenho racional de HMMs de perfil. Partindo de um alinhamento de múltiplas sequências, o TABAJARA é capaz de encontrar blocos que são (1) conservados ou (2) discriminativos para dois ou mais grupos de sequências. O programa utiliza diferentes métricas para atribuir pontuações posição-específicas ao longo de todo o alinhamento e utiliza então uma janela deslizante para encontrar as regiões com maiores pontuações. Blocos de alinhamento selecionados são então extraídos e utilizados para construir HMMs de perfil. Para validar o método, o programa TABAJARA foi empregado para a construção de modelos para vírus do gênero Flavivirus e para fagos da família Microviridae. Em ambos os grupos virais foi possível se obter modelos de ampla abrangência, capazes de detectar todos os membros de um respectivo grupo taxonômico, e modelos de abrangência mais restrita, específicos para espécies distintas de Flavivirus (ex. DENV, ZIKV ou YFV) ou subfamílias de Microviridae (ex. Alpavirinae, Gokushovirinae e Pichovirinae). Em outra validação, foram utilizadas sequências da endonuclease Cas1 para se obter modelos capazes de diferenciar CRISPRs de casposons, esses últimos representando uma superfamília de transposons de DNA autossintetizantes, os quais originaram o sistema de imunidade CRISPR-Cas de procariotos. O TABAJARA conseguiu gerar modelos específicos de Cas1 derivada de casposons, permitindo sua diferenciação em relação aos seus ortólogos de CRISPRs. No presente trabalho foi desenvolvido ainda o HMM-Prospector, uma ferramenta que utiliza um conjunto de HMMs de perfil para a triagem de dados de sequenciamento genômico ou metagenômico. O programa informa quais são os modelos mais reconhecidos pelas leituras, sob valores de corte de pontuação definidos pelo usuário, assim como quantas leituras são detectadas por cada modelo. Com esta informação, os modelos mais relevantes podem ser utilizados como sementes em montagens progressivas com o programa GenSeed-HMM, dentro de uma abordagem integrada para a construção de modelos e sua aplicação. Finamente, foi desenvolvido o e-Finder, um aplicativo genérico para a detecção e extração de elementos multigênicos a partir de genomas ou metagenomas montados utilizando HMMs de perfil. O e-Finder executa buscas de similaridade entre os HMMs de perfil e as sequências traduzidas dos dados montados e checa, em seguida, se os critérios de sintenia pré-definidos foram atendidos, incluindo o número mínimo de genes, a ordem dos genes e as distâncias intergênicas. As sequências dos elementos são então extraídas, as regiões codificantes (ORFs) identificadas e traduzidas conceitualmente em sequências completas de proteínas. Para validar esta ferramenta, foram empegados dois estudos de caso, profagos da família Microviridae e casposons, utilizando-se HMMs de perfil específicos, construídos com o programa TABAJARA. Em ambos os casos, o e-Finder foi executado usando-se a base de dados PATRIC, um repositório com mais de 135.000 genomas de bactérias e arqueias. Foram identificados um total de 91 contigs positivos para casposons a partir de 79 genomas distintos. No caso dos Microviridae, foram encontrados 104 profagos candidatos, estendendo o conhecimento da gama de hospedeiros bacterianos. Em ambos os casos, análises filogenéticas confirmaram a correta atribuição taxonômica das sequências positivas. Os programas desenvolvidos neste trabalho podem ser utilizados isoladamente ou em combinação para detectar e discriminar sequências conhecidas ou remotamente relacionadas. Juntamente com o GenSeed-HMM, estes programas constituem um conjunto integrado de ferramentas com potencial aplicação na busca de novos vírus e elementos genéticos móveis, bem como em qualquer outra tarefa relacionada à detecção e/ou discriminação de subgrupos de famílias de sequências nucleotídicas ou proteicas / Profile HMMs are a powerful way of modeling sequence diversity and constitute a very sensitive approach to detect remote orthologs. A potential application of such models is the detection of emerging viruses and novel mobile genetic elements. Our group has recently developed GenSeed-HMM, a tool that employs profile HMMs as seeds for gene-targeted progressive assembly using either genomic or metagenomic data. In this work we developed TABAJARA, a program for the rational design of profile HMMs. Starting from a multiple sequence alignment, TABAJARA is able to find blocks that are either (1) conserved across all sequences or (2) discriminative for two or more specific groups of sequences. The program uses different metrics to ascribe position-specific scores along the whole alignment and then uses a sliding-window to find top-scoring regions. Selected alignment blocks are then extracted and used to build profile HMMs. To validate the method, we employed TABAJARA to construct models for viruses of the Flavivirus genus and phages of the Microviridae family. In both viral groups we were able to obtain wide-range models, able to detect all members of the respective taxonomic group, and models that are specific to particular Flavivirus species (e.g. DENV, ZIKV or YFV) or Microviridae subfamilies (e.g. Alpavirinae, Gokushovirinae and Pichovirinae). In another validation, we used sequences of the endonuclease Cas1 to obtain models capable of differentiating CRISPRs from casposons, the latter elements representing a superfamily of self-synthesizing DNA transposons that originated the prokaryotic CRISPR-Cas immunity. TABAJARA succeeded to generate models specific to casposon-derived Cas1, enabling their differentiation from CRISPR orthologs. We also developed HMM-Prospector, a tool that can use a batch of profile HMMs to screen genomic or metagenomic sequencing data, reporting which profile HMMs are mostly recognized under user-defined score cutoff values, and how many reads are detected by each model. With this information, the most relevant models can be used as seeds in progressive assemblies with GenSeed-HMM program, providing an integrated approach for model construction and application. Finally, we developed e-Finder, a generic application for detecting and extracting multigene elements from assembled genomes or metagenomes using profile HMMs. e-Finder runs similarity searches of profile HMMs against translated sequences of the assembled data and then checks if pre-defined syntenic criteria have been fulfilled, including minimum number of genes, gene order and intergenic distances. Element sequences are then extracted, their ORFs identified and conceptually translated into full-length protein sequences. To validate the tool, we employed two distinct case studies, prophages of the Microviridae family and casposons, using specific profile HMMs constructed by TABAJARA. In both cases, we executed e-Finder using the PATRIC database, a repository with over 135,000 bacterial and archaeal genomes. We identified in total 91 casposon-positive contigs from 79 distinct genomes. In the case of Microviridae, we found a total of 104 provirus candidates, extending the known range of bacterial hosts. In both cases, phylogenetic analyses confirmed the correct taxonomic assignment of the positive sequences. The programs developed in this work can be used alone or in combination to detect and discriminate known or distantly related sequences. Together with GenSeed-HMM, these programs provide an integrated toolbox with potential application in the search of novel viruses and mobile genetic elements, as well as in any other task related to the detection and/or discrimination of subgroups of DNA or protein sequences. Detecção de vírus Famílias proteicas Hidden Markov models HMMs de perfil Metagenômica Metagenomics Modelos ocultos de Markov Profile HMMs Protein families Sintenia Synteny Virus detection
5	Chereme- Based Recognition of Isolated, Dynamic Gestures from South African Sign Language with Hidden Markov Models Rajah, Christopher January 2006 (has links) Masters of Science / Much work has been done in building systems that can recognise gestures, e.g. as a component of sign language recognition systems. These systems typically use whole gestures as the smallest unit for recognition. Although high recognition rates have been reported, these systems do not scale well and are computationally intensive. The reason why these systems generally scale poorly is that they recognize gestures by building individual models for each separate gesture; as the number of gestures grows, so does the required number of models. Beyond a certain threshold number of gestures to be recognized, this approach becomes infeasible. This work proposes that similarly good recognition rates can be achieved by building models for subcomponents of whole gestures, so-called cheremes. Instead of building models for entire gestures, we build models for cheremes and recognize gestures as sequences of such cheremes. The assumption is that many gestures share cheremes and that the number of cheremes necessary to describe gestures is much smaller than the number of gestures. This small number of cheremes then makes it possible to recognize a large number of gestures with a small number of chereme models. This approach is akin to phoneme-based speech recognition systems where utterances are recognized as phonemes which in turn are combined into words. We attempt to recognise and classify cheremes found in South African Sign Language (SASL). We introduce a method for the automatic discovery of cheremes in dynamic signs. We design, train and use hidden Markov models (HMMs) for chereme recognition. Our results show that this approach is feasible in that it not only scales well, but it also generalizes well. We are able to recognize cheremes in signs that were not used for training HMMs; this generalization ability is a basic necessity for chemere-based gesture recognition. Our approach can thus lay the foundation for building a SASL dynamic gesture recognition system. Cheremes Phonemes South African Sign Language (SASL) Hidden Markov models (HMMs)
6	Reconnaissance de caractères par méthodes markoviennes et réseaux bayésiens Hallouli, Khalid 05 1900 (has links) (PDF) Cette thése porte sur la reconnaissance de caractères imprimés et manuscrits par méthodes markoviennes et réseaux bayésiens. La première partie consiste à effectuer une modélisation stochastique markovienne en utilisant les HMMs classiques dans deux cas: semi-continu et discret. Un premier modèle HMM est obtenu à partir d'observations de type colonnes de pixels (HMM-vertical), le second à partir d'observations de type lignes (HMM-horizontal). Ensuite nous proposons deux types de modèles de fusion : modèle de fusion de scores qui consiste à combiner les deux vraisemblances résultantes des deux HMMs, et modèle de fusion de données qui regroupe simultanément les deux observations lignes et colonnes. Les résultats montrent l'importance du cas semi-continu et la performance des modèles de fusion. Dans la deuxième partie nous développons les réseaux bayésiens statiques et dynamiques, l'algorithme de Jensen Lauritzen Olesen (JLO) servant comme moteur d'inférence exacte, ainsi que l'apprentissage des paramètres avec des données complètes et incomplètes. Nous proposons une approche pour la reconnaissance de caractères (imprimés et manuscrits) en employant le formalisme des réseaux bayésiens dynamiques. Nous construisons certains types de modèles: HMM sous forme de réseau bayésien dynamique, modèle de trajectoire et modèles de couplages. Les résultats obtenus mettent en évidence la bonne performance des modèles couplés. En général nos applications nous permettent de conclure que l'utilisation des réseaux bayésiens est efficace et très prometteuse par le fait de modéliser les dépendances entre différentes observations dans les images de caractères. HMMs Réseau Bayésien statique et dynamique Inference Apprentissage Arbre de jonction Algorithme EM Quntification vectorielle Fusion
7	From protein sequence to structural instability and disease Wang, Lixiao January 2010 (has links) A great challenge in bioinformatics is to accurately predict protein structure and function from its amino acid sequence, including annotation of protein domains, identification of protein disordered regions and detecting protein stability changes resulting from amino acid mutations. The combination of bioinformatics, genomics and proteomics becomes essential for the investigation of biological, cellular and molecular aspects of disease, and therefore can greatly contribute to the understanding of protein structures and facilitating drug discovery. In this thesis, a PREDICTOR, which consists of three machine learning methods applied to three different but related structure bioinformatics tasks, is presented: using profile Hidden Markov Models (HMMs) to identify remote sequence homologues, on the basis of protein domains; predicting order and disorder in proteins using Conditional Random Fields (CRFs); applying Support Vector Machines (SVMs) to detect protein stability changes due to single mutation. To facilitate structural instability and disease studies, these methods are implemented in three web servers: FISH, OnD-CRF and ProSMS, respectively. For FISH, most of the work presented in the thesis focuses on the design and construction of the web-server. The server is based on a collection of structure-anchored hidden Markov models (saHMM), which are used to identify structural similarity on the protein domain level. For the order and disorder prediction server, OnD-CRF, I implemented two schemes to alleviate the imbalance problem between ordered and disordered amino acids in the training dataset. One uses pruning of the protein sequence in order to obtain a balanced training dataset. The other tries to find the optimal p-value cut-off for discriminating between ordered and disordered amino acids. Both these schemes enhance the sensitivity of detecting disordered amino acids in proteins. In addition, the output from the OnD-CRF web server can also be used to identify flexible regions, as well as predicting the effect of mutations on protein stability. For ProSMS, we propose, after careful evaluation with different methods, a clustered by homology and a non-clustered model for a three-state classification of protein stability changes due to single amino acid mutations. Results for the non-clustered model reveal that the sequence-only based prediction accuracy is comparable to the accuracy based on protein 3D structure information. In the case of the clustered model, however, the prediction accuracy is significantly improved when protein tertiary structure information, in form of local environmental conditions, is included. Comparing the prediction accuracies for the two models indicates that the prediction of mutation stability of proteins that are not homologous is still a challenging task. Benchmarking results show that, as stand-alone programs, these predictors can be comparable or superior to previously established predictors. Combined into a program package, these mutually complementary predictors will facilitate the understanding of structural instability and disease from protein sequence. protein domain remote homologue protein function point mutation protein family protein stability HMMs CRFs SVMs
8	Arabic Text Recognition and Machine Translation Alkhoury, Ihab 13 July 2015 (has links) [EN] Research on Arabic Handwritten Text Recognition (HTR) and Arabic-English Machine Translation (MT) has been usually approached as two independent areas of study. However, the idea of creating one system that combines both areas together, in order to generate English translation out of images containing Arabic text, is still a very challenging task. This process can be interpreted as the translation of Arabic images. In this thesis, we propose a system that recognizes Arabic handwritten text images, and translates the recognized text into English. This system is built from the combination of an HTR system and an MT system. Regarding the HTR system, our work focuses on the use of Bernoulli Hidden Markov Models (BHMMs). BHMMs had proven to work very well with Latin script. Indeed, empirical results based on it were reported on well-known corpora, such as IAM and RIMES. In this thesis, these results are extended to Arabic script, in particular, to the well-known IfN/ENIT and NIST OpenHaRT databases for Arabic handwritten text. The need for transcribing Arabic text is not only limited to handwritten text, but also to printed text. Arabic printed text might be considered as a simple form of handwritten text version. Thus, for this kind of text, we also propose Bernoulli HMMs. In addition, we propose to compare BHMMs with state-of-the-art technology based on neural networks. A key idea that has proven to be very effective in this application of Bernoulli HMMs is the use of a sliding window of adequate width for feature extraction. This idea has allowed us to obtain very competitive results in the recognition of both Arabic handwriting and printed text. Indeed, a system based on it ranked first at the ICDAR 2011 Arabic recognition competition on the Arabic Printed Text Image (APTI) database. Moreover, this idea has been refined by using repositioning techniques for extracted windows, leading to further improvements in Arabic text recognition. In the case of handwritten text, this refinement improved our system which ranked first at the ICFHR 2010 Arabic handwriting recognition competition on IfN/ENIT. In the case of printed text, this refinement led to an improved system which ranked second at the ICDAR 2013 Competition on Multi-font and Multi-size Digitally Represented Arabic Text on APTI. Furthermore, this refinement was used with neural networks-based technology, which led to state-of-the-art results. For machine translation, the system was based on the combination of three state-of-the-art statistical models: the standard phrase-based models, the hierarchical phrase-based models, and the N-gram phrase-based models. This combination was done using the Recognizer Output Voting Error Reduction (ROVER) method. Finally, we propose three methods of combining HTR and MT to develop an Arabic image translation system. The system was evaluated on the NIST OpenHaRT database, where competitive results were obtained. / [ES] El reconocimiento de texto manuscrito (HTR) en árabe y la traducción automática (MT) del árabe al inglés se han tratado habitualmente como dos áreas de estudio independientes. De hecho, la idea de crear un sistema que combine las dos áreas, que directamente genere texto en inglés a partir de imágenes que contienen texto en árabe, sigue siendo una tarea difícil. Este proceso se puede interpretar como la traducción de imágenes de texto en árabe. En esta tesis, se propone un sistema que reconoce las imágenes de texto manuscrito en árabe, y que traduce el texto reconocido al inglés. Este sistema está construido a partir de la combinación de un sistema HTR y un sistema MT. En cuanto al sistema HTR, nuestro trabajo se enfoca en el uso de los Bernoulli Hidden Markov Models (BHMMs). Los modelos BHMMs ya han sido probados anteriormente en tareas con alfabeto latino obteniendo buenos resultados. De hecho, existen resultados empíricos publicados usando corpus conocidos, tales como IAM o RIMES. En esta tesis, estos resultados se han extendido al texto manuscrito en árabe, en particular, a las bases de datos IfN/ENIT y NIST OpenHaRT. En aplicaciones reales, la transcripción del texto en árabe no se limita únicamente al texto manuscrito, sino también al texto impreso. El texto impreso se puede interpretar como una forma simplificada de texto manuscrito. Por lo tanto, para este tipo de texto, también proponemos el uso de modelos BHMMs. Además, estos modelos se han comparado con tecnología del estado del arte basada en redes neuronales. Una idea clave que ha demostrado ser muy eficaz en la aplicación de modelos BHMMs es el uso de una ventana deslizante (sliding window) de anchura adecuada durante la extracción de características. Esta idea ha permitido obtener resultados muy competitivos tanto en el reconocimiento de texto manuscrito en árabe como en el de texto impreso. De hecho, un sistema basado en este tipo de extracción de características quedó en la primera posición en el concurso ICDAR 2011 Arabic recognition competition usando la base de datos Arabic Printed Text Image (APTI). Además, esta idea se ha perfeccionado mediante el uso de técnicas de reposicionamiento aplicadas a las ventanas extraídas, dando lugar a nuevas mejoras en el reconocimiento de texto árabe. En el caso de texto manuscrito, este refinamiento ha conseguido mejorar el sistema que ocupó el primer lugar en el concurso ICFHR 2010 Arabic handwriting recognition competition usando IfN/ENIT. En el caso del texto impreso, este refinamiento condujo a un sistema mejor que ocupó el segundo lugar en el concurso ICDAR 2013 Competition on Multi-font and Multi-size Digitally Represented Arabic Text en el que se usaba APTI. Por otro lado, esta técnica se ha evaluado también en tecnología basada en redes neuronales, lo que ha llevado a resultados del estado del arte. Respecto a la traducción automática, el sistema se ha basado en la combinación de tres tipos de modelos estadísticos del estado del arte: los modelos standard phrase-based, los modelos hierarchical phrase-based y los modelos N-gram phrase-based. Esta combinación se hizo utilizando el método Recognizer Output Voting Error Reduction (ROVER). Por último, se han propuesto tres métodos para combinar los sistemas HTR y MT con el fin de desarrollar un sistema de traducción de imágenes de texto árabe a inglés. El sistema se ha evaluado sobre la base de datos NIST OpenHaRT, donde se han obtenido resultados competitivos. / [CA] El reconeixement de text manuscrit (HTR) en àrab i la traducció automàtica (MT) de l'àrab a l'anglès s'han tractat habitualment com dues àrees d'estudi independents. De fet, la idea de crear un sistema que combine les dues àrees, que directament genere text en anglès a partir d'imatges que contenen text en àrab, continua sent una tasca difícil. Aquest procés es pot interpretar com la traducció d'imatges de text en àrab. En aquesta tesi, es proposa un sistema que reconeix les imatges de text manuscrit en àrab, i que tradueix el text reconegut a l'anglès. Aquest sistema està construït a partir de la combinació d'un sistema HTR i d'un sistema MT. Pel que fa al sistema HTR, el nostre treball s'enfoca en l'ús dels Bernoulli Hidden Markov Models (BHMMs). Els models BHMMs ja han estat provats anteriorment en tasques amb alfabet llatí obtenint bons resultats. De fet, existeixen resultats empírics publicats emprant corpus coneguts, tals com IAM o RIMES. En aquesta tesi, aquests resultats s'han estès a la escriptura manuscrita en àrab, en particular, a les bases de dades IfN/ENIT i NIST OpenHaRT. En aplicacions reals, la transcripció de text en àrab no es limita únicament al text manuscrit, sinó també al text imprès. El text imprès es pot interpretar com una forma simplificada de text manuscrit. Per tant, per a aquest tipus de text, també proposem l'ús de models BHMMs. A més a més, aquests models s'han comparat amb tecnologia de l'estat de l'art basada en xarxes neuronals. Una idea clau que ha demostrat ser molt eficaç en l'aplicació de models BHMMs és l'ús d'una finestra lliscant (sliding window) d'amplària adequada durant l'extracció de característiques. Aquesta idea ha permès obtenir resultats molt competitius tant en el reconeixement de text àrab manuscrit com en el de text imprès. De fet, un sistema basat en aquest tipus d'extracció de característiques va quedar en primera posició en el concurs ICDAR 2011 Arabic recognition competition emprant la base de dades Arabic Printed Text Image (APTI). A més a més, aquesta idea s'ha perfeccionat mitjançant l'ús de tècniques de reposicionament aplicades a les finestres extretes, donant lloc a noves millores en el reconeixement de text en àrab. En el cas de text manuscrit, aquest refinament ha aconseguit millorar el sistema que va ocupar el primer lloc en el concurs ICFHR 2010 Arabic handwriting recognition competition usant IfN/ENIT. En el cas del text imprès, aquest refinament va conduir a un sistema millor que va ocupar el segon lloc en el concurs ICDAR 2013 Competition on Multi-font and Multi-size Digitally Represented Arabic Text en el qual s'usava APTI. D'altra banda, aquesta tècnica s'ha avaluat també en tecnologia basada en xarxes neuronals, el que ha portat a resultats de l'estat de l'art. Respecte a la traducció automàtica, el sistema s'ha basat en la combinació de tres tipus de models estadístics de l'estat de l'art: els models standard phrase-based, els models hierarchical phrase-based i els models N-gram phrase-based. Aquesta combinació es va fer utilitzant el mètode Recognizer Output Voting Errada Reduction (ROVER). Finalment, s'han proposat tres mètodes per combinar els sistemes HTR i MT amb la finalitat de desenvolupar un sistema de traducció d'imatges de text àrab a anglès. El sistema s'ha avaluat sobre la base de dades NIST OpenHaRT, on s'han obtingut resultats competitius. / Alkhoury, I. (2015). Arabic Text Recognition and Machine Translation [Tesis doctoral]. Universitat Politècnica de València. https://doi.org/10.4995/Thesis/10251/53029 Arabic Image Translation Arabic OCR Arabic Machine Translation Arabic text recognition Bernoulli HMMs ESTADISTICA E INVESTIGACION OPERATIVA LENGUAJES Y SISTEMAS INFORMATICOS
9	Construção e aplicação de HMMs de perfil para a detecção e classificação de vírus / Construction and application of profile HMMs for the specific detection and classification of viruses Guimarães, Miriã Nunes 22 February 2019 (has links) Os vírus são as entidades biológicas mais abundantes encontradas na natureza. O método clássico de estudo dos vírus requerem seu isolamento e propagação in vitro. Contudo, necessita-se ter um conhecimento prévio sobre as condições necessárias para seu cultivo em células, sendo assim a maior parte dos vírus existentes não é conhecida. Análises metagenômicas são uma alternativa para a detecção e caracterização de novos vírus, uma vez que não requerem um cultivo prévio e as amostras podem conter material genético de múltiplos organismos. Uma vez obtidas as sequências montadas a partir das leituras metagenômicas, o método mais utilizado para a identificação e classificação dos organismos é a busca de similaridade com o programa BLAST contra bancos de sequências conhecidas. Contudo, métodos de alinhamento pareado são capazes de identificar apenas sequências com identidade superior a 20-30%. Uma alternativa a essa limitação é o uso de métodos baseados no uso de perfis, que podem aumentar a sensibilidade de detecção de homólogos filogeneticamente distantes. HMMs de perfil são modelos probabilísticos capazes de representar a diversidade de caracteres em posições-específicas de um alinhamento de múltiplas sequências. Nosso grupo desenvolveu a ferramenta TABAJARA, utilizada neste projeto, para a identificação de blocos que podem ser conservados em todas as sequências do alinhamento ou discriminativos entre grupos de sequências. Esses blocos são utilizados para a geração de HMMs de perfil, os quais podem ser usados, no contexto da virologia, para a identificação de grupos taxonômicos amplos como famílias virais ou, ainda, taxa mais restritos como gêneros ou mesmo espécies de vírus. O presente projeto teve como objetivos aplicar e otimizar o programa TABAJARA em diferentes grupos taxonômicos de vírus, construir modelos específicos para cada um desses grupos e validar esses modelos em dados metagenômicos. O primeiro modelo de estudo escolhido foi a ordem Bunyavirales, composta de vírus de ssRNA (-) majoritariamente envelopados e esféricos, com genoma segmentado e pertencentes ao grupo 5 da classificação de Baltimore. Este grupo inclui vírus causadores de várias doenças em humanos, animais e plantas. O segundo modelo de estudo escolhido foi a família Togaviridae, composta de vírus de ssRNA (+) envelopados e esféricos, cujo genoma expressa uma poliproteína e pertencem ao grupo 4 da classificação de Baltimore. Este grupo inclui o vírus Chikungunya e outras espécies que causam diversas patologias ao homem. O terceiro modelo de estudo escolhido foi a subfamília Spounavirinae, compreendendo bacteriófagos que infectam vários hospedeiros bacterianos e em alguns casos possuem potencial terapêutico comprovado contra infecções bacterianas que afetam o homem. Estes fagos apresentam partículas virais com estrutura cabeça-cauda, não são envelopados, apresentam genoma de dsDNA e pertencem ao grupo 1 da classificação de Baltimore. Todos os modelos construídos foram validados quanto à sensibilidade e especificidade de detecção e, ao final, foram utilizados em análises de prospecção de vírus em dados metagenômicos obtidos na base SRA do NCBI. Os HMMs de perfil apresentaram excelente desempenho, comprovando a viabilidade da metodologia proposta neste projeto. Os resultados apresentados neste trabalho abrem a perspectiva da ampla utilização de HMMs de perfil como ferramentas universais para a detecção e classificação de vírus em dados metagenômicos. / Viruses are the most widely biological entities found in nature. Most of the information that can be obtained from these organisms requires viral in vitro isolation and cultivation. However, most of the existing viruses are still unknown because the biological requirements for their successful propagation have not been identified so far. Metagenomic analyses offer an interesting alternative for the detection and characterization of novel viruses, since previous cultivation is not required, and the samples may contain genetic material of multiple organisms. Once assembled sequences are obtained from individual reads, the most widely used method for viral identification and classification is the use of BLAST similarity searches against databases of known sequences. However, pairwise alignment methods are only able to identify sequences that present identity greater than 20-30%. Profile-based methods may increase the sensitivity of detection of remote homologues. Profile HMMs are probabilistic models capable of representing the diversity of amino acid residues at specific positions of a multiple sequence alignment. Our group is developing TABAJARA, a tool for the identification of alignment blocks that are conserved across all sequences of the alignment or discriminative between groups of sequences. These blocks are used to generate profile HMMs, which can in turn be used, in the context of virology, to identify broad taxonomic groups, such as viral families, or narrower taxa as genera or viral species. The present project aimed to apply and standardize the use of TABAJARA in different taxonomic groups of viruses, to build specific models for each of these groups and to validate these models in metagenomic data. We used three viral models for this study. The first chosen model was the Bunyavirales order, composed of mostly enveloped and spherical ssRNA(-) viruses with a segmented genome belonging to group 5 of the Baltimore classification. This group includes viruses that cause several important diseases in humans, animals and plants. The second chosen model was the Togaviridae family, composed of enveloped and spherical ssRNA(+) viruses, with a genome coding for a polyprotein, and belonging to group 4 of the Baltimore classification. This group includes the Chikungunya virus and some other viral species that cause relevant pathologies to humans and animals. Finally, we used the Spounavirinae subfamily, comprising viruses that infect a variety of bacterial hosts and that can potentially be used for phage therapy of some human bacterial diseases. These phages present non-enveloped virions with a head-to-tail structure, a dsDNA genome, and belong to group 1 of the Baltimore classification. All constructed profile HMMs were evaluated in regard to their sensitivity and specificity of detection, as well as tested in viral surveys using metagenomic data from the SRA database. The profile HMMs presented excellent performance, proving the viability of the methodology proposed in this project. The results presented in this work open the perspective of the wide use of profile HMMs as universal tools for the detection and classification of viruses in metagenomic data. Bioinformática Genomas Hidden Markov models Marcador molecular Metagenomics Modelos para processos estocásticos Molecular markers Profile HMMs Viral taxonomy Vírus Virus detection
10	Speech Signal Classification Using Support Vector Machines Sood, Gaurav 07 1900 (has links) Hidden Markov Models (HMMs) are, undoubtedly, the most employed core technique for Automatic Speech Recognition (ASR). Nevertheless, we are still far from achieving high‐performance ASR systems. Some alternative approaches, most of them based on Artificial Neural Networks (ANNs), were proposed during the late eighties and early nineties. Some of them tackled the ASR problem using predictive ANNs, while others proposed hybrid HMM/ANN systems. However, despite some achievements, nowadays, the dependency on Hidden Markov Models is a fact. During the last decade, however, a new tool appeared in the field of machine learning that has proved to be able to cope with hard classification problems in several fields of application: the Support Vector Machines (SVMs). The SVMs are effective discriminative classifiers with several outstanding characteristics, namely: their solution is that with maximum margin; they are capable to deal with samples of a very higher dimensionality; and their convergence to the minimum of the associated cost function is guaranteed. In this work a novel approach based upon probabilistic kernels in support vector machines have been attempted for speech data classification. The classification accuracy in case of support vector classification depends upon the kernel function used which in turn depends upon the data set in hand. But still as of now there is no way to know a priori which kernel will give us best results The kernel used in this work tries to normalize the time dimension by fitting a probability distribution over individual data points which normalizes the time dimension inherent to speech signals which facilitates the use of support vector machines since it acts on static data only. The divergence between these probability distributions fitted over individual speech utterances is used to form the kernel matrix. Vowel Classification, Isolated Word Recognition (Digit Recognition), have been attempted and results are compared with state of art systems. Speech Recognition Speech Signal Processing Automatic Speech Recognition Artificial Neural Networks Support Vector Machine Time Normalization Hidden Markov Models (HMMs) Computer Science

Search results