LAUSR.org creates dashboard-style pages of related content for over 1.5 million academic articles. Sign Up to like articles & get recommendations!

Leveraging Large Language Models for Dataset Discovery

The exponential growth of data across diverse domains highlights the need for efficient methods in discovering relevant datasets. Traditional search engines such as Google have served as the go‐to tools… Click to show full abstract

The exponential growth of data across diverse domains highlights the need for efficient methods in discovering relevant datasets. Traditional search engines such as Google have served as the go‐to tools for this purpose. Recent advancements in large language models (LLMs) such as ChatGPT and Microsoft Copilot have sparked interest in their potential to serve as alternatives for data discovery. While these models are primarily designed for conversational interactions, their capabilities in information retrieval and dataset discovery are becoming areas of active exploration. In this work, we present a mixed‐method study that investigates the difference in user experience when using Google and Microsoft Copilot to search for datasets. This study aims to uncover the strengths and limitations of LLMs in data discovery, offering insights into their potential as alternatives or complements to traditional tools.

Keywords: large language; language; leveraging large; language models; discovery; dataset discovery

Journal Title: Proceedings of the Association for Information Science and Technology
Year Published: 2025

Link to full text (if available)


Share on Social Media:                               Sign Up to like & get
recommendations!

Related content

More Information              News              Social Media              Video              Recommended



                Click one of the above tabs to view related content.