Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Tommaso Soru

Leveraging Log Probabilities in Language Models to Forecast Future Events

Jan 08, 2025

Tommaso Soru, Jim Marshall

Figure 1 for Leveraging Log Probabilities in Language Models to Forecast Future Events

Figure 2 for Leveraging Log Probabilities in Language Models to Forecast Future Events

Figure 3 for Leveraging Log Probabilities in Language Models to Forecast Future Events

Figure 4 for Leveraging Log Probabilities in Language Models to Forecast Future Events

Abstract:In the constantly changing field of data-driven decision making, accurately predicting future events is crucial for strategic planning in various sectors. The emergence of Large Language Models (LLMs) marks a significant advancement in this area, offering advanced tools that utilise extensive text data for prediction. In this industry paper, we introduce a novel method for AI-driven foresight using LLMs. Building on top of previous research, we employ data on current trends and their trajectories for generating forecasts on 15 different topics. Subsequently, we estimate their probabilities via a multi-step approach based on log probabilities. We show we achieve a Brier score of 0.186, meaning a +26% improvement over random chance and a +19% improvement over widely-available AI systems.

* 5 pages, 4 figures

Via

Access Paper or Ask Questions

Exploring Sequence-to-Sequence Models for SPARQL Pattern Composition

Oct 21, 2020

Anand Panchbhai, Tommaso Soru, Edgard Marx

Figure 1 for Exploring Sequence-to-Sequence Models for SPARQL Pattern Composition

Figure 2 for Exploring Sequence-to-Sequence Models for SPARQL Pattern Composition

Abstract:A booming amount of information is continuously added to the Internet as structured and unstructured data, feeding knowledge bases such as DBpedia and Wikidata with billions of statements describing millions of entities. The aim of Question Answering systems is to allow lay users to access such data using natural language without needing to write formal queries. However, users often submit questions that are complex and require a certain level of abstraction and reasoning to decompose them into basic graph patterns. In this short paper, we explore the use of architectures based on Neural Machine Translation called Neural SPARQL Machines to learn pattern compositions. We show that sequence-to-sequence models are a viable and promising option to transform long utterances into complex SPARQL queries.

* Proceedings of the First Indo-American Knowledge Graph and Semantic Web Conference (KGSWC-India 2020)

Via

Access Paper or Ask Questions

Where is Linked Data in Question Answering over Linked Data?

May 07, 2020

Tommaso Soru, Edgard Marx, André Valdestilhas, Diego Moussallem, Gustavo Publio, Muhammad Saleem

Figure 1 for Where is Linked Data in Question Answering over Linked Data?

Abstract:We argue that "Question Answering with Knowledge Base" and "Question Answering over Linked Data" are currently two instances of the same problem, despite one explicitly declares to deal with Linked Data. We point out the lack of existing methods to evaluate question answering on datasets which exploit external links to the rest of the cloud or share common schema. To this end, we propose the creation of new evaluation settings to leverage the advantages of the Semantic Web to achieve AI-complete question answering.

* Position paper, THE Workshop @ ISWC 2018

Via

Access Paper or Ask Questions

Concept2vec: Metrics for Evaluating Quality of Embeddings for Ontological Concepts

Jul 26, 2018

Faisal Alshargi, Saeedeh Shekarpour, Tommaso Soru, Amit Sheth

Figure 1 for Concept2vec: Metrics for Evaluating Quality of Embeddings for Ontological Concepts

Figure 2 for Concept2vec: Metrics for Evaluating Quality of Embeddings for Ontological Concepts

Figure 3 for Concept2vec: Metrics for Evaluating Quality of Embeddings for Ontological Concepts

Figure 4 for Concept2vec: Metrics for Evaluating Quality of Embeddings for Ontological Concepts

Abstract:Although there is an emerging trend towards generating embeddings for primarily unstructured data, and recently for structured data, there is not yet any systematic suite for measuring the quality of embeddings. This deficiency is further sensed with respect to embeddings generated for structured data because there are no concrete evaluation metrics measuring the quality of encoded structure as well as semantic patterns in the embedding space. In this paper, we introduce a framework containing three distinct tasks concerned with the individual aspects of ontological concepts: (i) the categorization aspect, (ii) the hierarchical aspect, and (iii) the relational aspect. Then, in the scope of each task, a number of intrinsic metrics are proposed for evaluating the quality of the embeddings. Furthermore, w.r.t. this framework multiple experimental studies were run to compare the quality of the available embedding models. Employing this framework in future research can reduce misjudgment and provide greater insight about quality comparisons of embeddings for ontological concepts.

* Working paper

Via

Access Paper or Ask Questions

ML-Schema: Exposing the Semantics of Machine Learning with Schemas and Ontologies

Jul 14, 2018

Gustavo Correa Publio, Diego Esteves, Agnieszka Ławrynowicz, Panče Panov, Larisa Soldatova, Tommaso Soru, Joaquin Vanschoren, Hamid Zafar

Figure 1 for ML-Schema: Exposing the Semantics of Machine Learning with Schemas and Ontologies

Figure 2 for ML-Schema: Exposing the Semantics of Machine Learning with Schemas and Ontologies

Figure 3 for ML-Schema: Exposing the Semantics of Machine Learning with Schemas and Ontologies

Abstract:The ML-Schema, proposed by the W3C Machine Learning Schema Community Group, is a top-level ontology that provides a set of classes, properties, and restrictions for representing and interchanging information on machine learning algorithms, datasets, and experiments. It can be easily extended and specialized and it is also mapped to other more domain-specific ontologies developed in the area of machine learning and data mining. In this paper we overview existing state-of-the-art machine learning interchange formats and present the first release of ML-Schema, a canonical format resulted of more than seven years of experience among different research institutions. We argue that exposing semantics of machine learning algorithms, models, and experiments through a canonical format may pave the way to better interpretability and to realistically achieve the full interoperability of experiments regardless of platform or adopted workflow solution.

* Poster, selected for the 2nd Reproducibility in Machine Learning Workshop at ICML 2018, Stockholm, Sweden

Via

Access Paper or Ask Questions

Neural Machine Translation for Query Construction and Composition

Jul 09, 2018

Tommaso Soru, Edgard Marx, André Valdestilhas, Diego Esteves, Diego Moussallem, Gustavo Publio

Figure 1 for Neural Machine Translation for Query Construction and Composition

Figure 2 for Neural Machine Translation for Query Construction and Composition

Figure 3 for Neural Machine Translation for Query Construction and Composition

Abstract:Research on question answering with knowledge base has recently seen an increasing use of deep architectures. In this extended abstract, we study the application of the neural machine translation paradigm for question parsing. We employ a sequence-to-sequence model to learn graph patterns in the SPARQL graph query language and their compositions. Instead of inducing the programs through question-answer pairs, we expect a semi-supervised approach, where alignments between questions and queries are built through templates. We argue that the coverage of language utterances can be expanded using late notable works in natural language generation.

* ICML workshop on Neural Abstract Machines & Program Induction v2 (NAMPI), extended abstract

Via

Access Paper or Ask Questions

Expeditious Generation of Knowledge Graph Embeddings

Mar 21, 2018

Tommaso Soru, Stefano Ruberto, Diego Moussallem, Edgard Marx, Diego Esteves, Axel-Cyrille Ngonga Ngomo

Figure 1 for Expeditious Generation of Knowledge Graph Embeddings

Figure 2 for Expeditious Generation of Knowledge Graph Embeddings

Figure 3 for Expeditious Generation of Knowledge Graph Embeddings

Figure 4 for Expeditious Generation of Knowledge Graph Embeddings

Abstract:Knowledge Graph Embedding methods aim at representing entities and relations in a knowledge base as points or vectors in a continuous vector space. Several approaches using embeddings have shown promising results on tasks such as link prediction, entity recommendation, question answering, and triplet classification. However, only a few methods can compute low-dimensional embeddings of very large knowledge bases. In this paper, we propose KG2Vec, a novel approach to Knowledge Graph Embedding based on the skip-gram model. Instead of using a predefined scoring function, we learn it relying on Long Short-Term Memories. We evaluated the goodness of our embeddings on knowledge graph completion and show that KG2Vec is comparable to the quality of the scalable state-of-the-art approaches and can process large graphs by parsing more than a hundred million triples in less than 6 hours on common hardware.

* Submitted, 6 pages

Via

Access Paper or Ask Questions

Beyond Markov Logic: Efficient Mining of Prediction Rules in Large Graphs

Feb 13, 2018

Tommaso Soru, André Valdestilhas, Edgard Marx, Axel-Cyrille Ngonga Ngomo

Figure 1 for Beyond Markov Logic: Efficient Mining of Prediction Rules in Large Graphs

Figure 2 for Beyond Markov Logic: Efficient Mining of Prediction Rules in Large Graphs

Figure 3 for Beyond Markov Logic: Efficient Mining of Prediction Rules in Large Graphs

Figure 4 for Beyond Markov Logic: Efficient Mining of Prediction Rules in Large Graphs

Abstract:Graph representations of large knowledge bases may comprise billions of edges. Usually built upon human-generated ontologies, several knowledge bases do not feature declared ontological rules and are far from being complete. Current rule mining approaches rely on schemata or store the graph in-memory, which can be unfeasible for large graphs. In this paper, we introduce HornConcerto, an algorithm to discover Horn clauses in large graphs without the need of a schema. Using a standard fact-based confidence score, we can mine close Horn rules having an arbitrary body size. We show that our method can outperform existing approaches in terms of runtime and memory consumption and mine high-quality rules for the link prediction task, achieving state-of-the-art results on a widely-used benchmark. Moreover, we find that rules alone can perform inference significantly faster than embedding-based methods and achieve accuracies on link prediction comparable to resource-demanding approaches such as Markov Logic Networks.

* 13 pages, 4 figures

Via

Access Paper or Ask Questions

Mandolin: A Knowledge Discovery Framework for the Web of Data

Nov 03, 2017

Tommaso Soru, Diego Esteves, Edgard Marx, Axel-Cyrille Ngonga Ngomo

Figure 1 for Mandolin: A Knowledge Discovery Framework for the Web of Data

Figure 2 for Mandolin: A Knowledge Discovery Framework for the Web of Data

Figure 3 for Mandolin: A Knowledge Discovery Framework for the Web of Data

Figure 4 for Mandolin: A Knowledge Discovery Framework for the Web of Data

Abstract:Markov Logic Networks join probabilistic modeling with first-order logic and have been shown to integrate well with the Semantic Web foundations. While several approaches have been devised to tackle the subproblems of rule mining, grounding, and inference, no comprehensive workflow has been proposed so far. In this paper, we fill this gap by introducing a framework called Mandolin, which implements a workflow for knowledge discovery specifically on RDF datasets. Our framework imports knowledge from referenced graphs, creates similarity relationships among similar literals, and relies on state-of-the-art techniques for rule mining, grounding, and inference computation. We show that our best configuration scales well and achieves at least comparable results with respect to other statistical-relational-learning algorithms on link prediction.

* 6 pages

Via

Access Paper or Ask Questions

SPARQL as a Foreign Language

Aug 25, 2017

Tommaso Soru, Edgard Marx, Diego Moussallem, Gustavo Publio, André Valdestilhas, Diego Esteves, Ciro Baron Neto

Figure 1 for SPARQL as a Foreign Language

Figure 2 for SPARQL as a Foreign Language

Figure 3 for SPARQL as a Foreign Language

Figure 4 for SPARQL as a Foreign Language

Abstract:In the last years, the Linked Data Cloud has achieved a size of more than 100 billion facts pertaining to a multitude of domains. However, accessing this information has been significantly challenging for lay users. Approaches to problems such as Question Answering on Linked Data and Link Discovery have notably played a role in increasing information access. These approaches are often based on handcrafted and/or statistical models derived from data observation. Recently, Deep Learning architectures based on Neural Networks called seq2seq have shown to achieve state-of-the-art results at translating sequences into sequences. In this direction, we propose Neural SPARQL Machines, end-to-end deep architectures to translate any natural language expression into sentences encoding SPARQL queries. Our preliminary results, restricted on selected DBpedia classes, show that Neural SPARQL Machines are a promising approach for Question Answering on Linked Data, as they can deal with known problems such as vocabulary mismatch and perform graph pattern composition.

* SEMANTiCS 2017; 13th International Conference on Semantic Systems, 2017

Via

Access Paper or Ask Questions