Global ETD Search

1	Efficient Graph Summarization of Large Networks Hajiabadi, Mahdi 24 June 2022 (has links) In this thesis, we study the notion of graph summarization, which is a fundamental task of finding a compact representation of the original graph called the summary. Graph summarization can be used for reducing the footprint of the input graph, better visualization, anonymizing the identity of users, and query answering. There are two different frameworks of graph summarization we consider in this thesis, the utility-based framework and the correction set-based framework. In the utility-based framework, the input graph is summarized until a utility threshold is not violated. In the correction set-based framework a set of correction edges is produced along with the summary graph. In this thesis we propose two algorithms for the utility-based framework and one for the correction set-based framework. All these three algorithms are for static graphs (i.e. graphs that do not change over time). Then, we propose two more utility-based algorithms for fully dynamic graphs (i.e. graphs with edge insertions and deletions). Algorithms for graph summarization can be lossless (summarizing the input graph without losing any information) or lossy (losing some information about the input graph in order to summarize it more). Some of our algorithms are lossless and some lossy, but with controlled utility loss. Our first utility-driven graph summarization algorithm, G-SCIS, is based on a clique and independent set decomposition, that produces optimal compression with zero loss of utility. The compression provided is significantly better than state-of-the-art in lossless graph summarization, while the runtime is two orders of magnitude lower. Our second algorithm is T-BUDS, a highly scalable, utility-driven algorithm for fully controlled lossy summarization. It achieves high scalability by combining memory reduction using Maximum Spanning Tree with a novel binary search procedure. T-BUDS outperforms state-of-the-art drastically in terms of the quality of summarization and is about two orders of magnitude better in terms of speed. In contrast to the competition, we are able to handle web-scale graphs in a single machine without performance impediment as the utility threshold (and size of summary) decreases. Also, we show that our graph summaries can be used as-is to answer several important classes of queries, such as triangle enumeration, Pagerank and shortest paths. We then propose algorithm LDME, a correction set-based graph summarization algorithm that produces compact output representations in a fast and scalable manner. To achieve this, we introduce (1) weighted locality sensitive hashing to drastically reduce the number of comparisons required to find good node merges, (2) an efficient way to compute the best quality merges that produces more compact outputs, and (3) a new sort-based encoding algorithm that is faster and more robust. More interestingly, our algorithm provides performance tuning settings to allow the option of trading compression for running time. On high compression settings, LDME achieves compression equal to or better than the state of the art with up to 53x speedup in running time. On high speed settings, LDME achieves up to two orders of magnitude speedup with only slightly lower compression. We also present two lossless summarization algorithms, Optimal and Scalable, for summarizing fully dynamic graphs. More concretely, we follow the framework of G-SCIS, which produces summaries that can be used as-is in several graph analytics tasks. Different from G-SCIS, which is a batch algorithm, Optimal and Scalable are fully dynamic and can respond rapidly to each change in the graph. Not only are Optimal and Scalable able to outperform G-SCIS and other batch algorithms by several orders of magnitude, but they also significantly outperform MoSSo, the state-of-the-art in lossless dynamic graph summarization. While Optimal produces always the most optimal summary, Scalable is able to trade the amount of node reduction for extra scalability. For reasonable values of the parameter $K$, Scalable is able to outperform Optimal by an order of magnitude in speed, while keeping the rate of node reduction close to that of Optimal. An interesting fact that we observed experimentally is that even if we were to run a batch algorithm, such as G-SCIS, once for every big batch of changes, still they would be much slower than Scalable. For instance, if 1 million changes occur in a graph, Scalable is two orders of magnitude faster than running G-SCIS just once at the end of the 1 million-edge sequence. / Graduate Graph Summarization Query Answering Lossless summary Lossy summary Locality Sensitive Hashing Jaccard Similarity Weighted Jaccard Similarity Hashing Incremental Algorithms Randomized Algorithms
2	GENERATING SQL FROM NATURAL LANGUAGE IN FEW-SHOT AND ZERO-SHOT SCENARIOS Asplund, Liam January 2024 (has links) Making information stored in databases more accessible to users inexperienced in structured query language (SQL) by converting natural language to SQL queries has long been a prominent research area in both the database and natural language processing (NLP) communities. There have been numerous approaches proposed for this task, such as encoder-decoder frameworks, semantic grammars, and more recently with the use of large language models (LLMs). When training LLMs to successfully generate SQL queries from natural language questions there are three notable methods used, pretraining, transfer learning and in-context learning (ICL). ICL is particularly advantageous in scenarios where the hardware at hand is limited, time is of concern and large amounts of task specific labled data is nonexistent. This study seeks to evaluate two strategies in ICL, namely zero-shot and few-shot scenarios using the Mistral-7B-Instruct LLM. Evaluation of the few-shot scenarios was conducted using two techniques, random selection and Jaccard Similarity. The zero-shot scenarios served as a baseline for the few-shot scenarios to overcome, which ended as anticipated, with the few-shot scenarios using Jaccard similarity outperforming the other two methods, followed by few-shot scenarios using random selection coming in at second best, and the zero-shot scenarios performing the worst. Evaluation results acquired based on execution accuracy and exact matching accuracy confirm that leveraging similarity in demonstrating examples when prompting the LLM will enhance the models knowledge about the database schema and table names which is used during the inference phase leadning to more accurately generated SQL queries than leveraging diversity in demonstrating examples. In-context learning Few-shot scenarios Zero-shot scenarios Large language model Prompt engineering Jaccard Similarity Computer Sciences Datavetenskap (datalogi)
3	Identification of common and unique stress responsive genes of Arabidopsis thaliana under different abiotic stress through RNA-Seq meta-analysis Akter, Shamima 06 February 2018 (has links) Abiotic stress is a major constraint for crop productivity worldwide. To better understand the common biological mechanisms of abiotic stress responses in plants, we performed meta-analysis of 652 samples of RNA sequencing (RNA-Seq) data from 43 published abiotic stress experiments in Arabidopsis thaliana. These samples were categorized into eight different abiotic stresses including drought, heat, cold, salt, light and wounding. We developed a multi-step computational pipeline, which performs data downloading, preprocessing, read mapping, read counting and differential expression analyses for RNA-Seq data. We found that 5729 and 5062 genes are induced or repressed by only one type of abiotic stresses. There are only 18 and 12 genes that are induced or repressed by all stresses. The commonly induced genes are related to gene expression regulation by stress hormone abscisic acid. The commonly repressed genes are related to reduced growth and chloroplast activities. We compared stress responsive genes between any two types of stresses and found that heat and cold regulate similar set of genes. We also found that high light affects different set of genes than blue light and red light. Interestingly, ABA regulated genes are different from those regulated by other stresses. Finally, we found that membrane related genes are repressed by ABA, heat, cold and wounding but are up regulated by blue light and red light. The results from this work will be used to further characterize the gene regulatory networks underlying stress responsive genes in plants. / Master of Science / Abiotic stress is a major constraint for crop productivity worldwide. To better understand the common biological mechanisms of abiotic stress responses in plants, we performed analysis of 652 samples of RNA sequencing data from 43 published abiotic stress experiments in Arabidopsis thaliana. These samples were collected from eight different abiotic stresses including drought, heat, cold, salt, light and wounding. We identified genes that were induced or repressed by each of these stresses. We found that 5729 and 5062 genes are induced or repressed by only one type of abiotic stresses. There are only 18 and 12 genes that are induced or repressed by all stresses. The commonly induced genes are related to gene expression regulation by stress hormone. The commonly repressed genes are related to reduced growth. We compared stress responsive genes between any two types of stresses and found that heat and cold regulate similar set of genes. We also found that high light affects different set of genes than blue light and red light. Finally, we found that membrane related genes are repressed by stress hormone, heat, cold and wounding but are up regulated by blue light and red light. The results from this work will be used to further characterize the gene regulations underlying stress responsive genes in plants. Abiotic stress Reactive oxygen species RNA-Seq RNA-Seq pipeline Gene Omnibus Series (GSE) Differentially Expressed Genes (DEG) Jaccard similarity index Gene Ontology (GO)

1

Page generated in 0.0568 seconds