UQV100: A test collection with query variability

Bailey, P, Moffat, A, Scholer, F and Thomas, P 2016, 'UQV100: A test collection with query variability', in Proceedings of the 39th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2016), Pisa, Italy, 17 - 21 July 2016, pp. 725-728.


Document type: Conference Paper
Collection: Conference Papers

Title UQV100: A test collection with query variability
Author(s) Bailey, P
Moffat, A
Scholer, F
Thomas, P
Year 2016
Conference name SIGIR 2016
Conference location Pisa, Italy
Conference dates 17 - 21 July 2016
Proceedings title Proceedings of the 39th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2016)
Publisher ACM
Place of publication United States
Start page 725
End page 728
Total pages 4
Abstract We describe the UQV100 test collection, designed to incorporate variability from users. Information need "backstories" were written for 100 topics (or sub-topics) from the TREC 2013 and 2014 Web Tracks. Crowd workers were asked to read the backstories, and provide the queries they would use; plus effort estimates of how many useful documents they would have to read to satisfy the need. A total of 10,835 queries were collected from 263 workers. After normalization and spell-correction, 5,764 unique variations remained; these were then used to construct a document pool via Indri-BM25 over the ClueWeb12-B corpus. Qualified crowd workers made relevance judgments relative to the backstories, using a relevance scale similar to the original TREC approach; first to a pool depth of ten per query, then deeper on a set of targeted documents. The backstories, query variations, normalized and spell-corrected queries, effort estimates, run outputs, and relevance judgments are made available collectively as the UQV100 test collection. We also make available the judging guidelines and the gold hits we used for crowd-worker qualification and spam detection. We believe this test collection will unlock new opportunities for novel investigations and analysis, including for problems such as task-intent retrieval performance and consistency (independent of query variation), query clustering, query difficulty prediction, and relevance feedback, among others.
Subjects Information Retrieval and Web Search
Copyright notice © 2016 ACM
ISBN 9781450340694
Versions
Version Filter Type
Citation counts: Scopus Citation Count Cited 6 times in Scopus Article | Citations
Access Statistics: 266 Abstract Views  -  Detailed Statistics
Created: Tue, 21 Mar 2017, 13:00:00 EST by Catalyst Administrator
© 2014 RMIT Research Repository • Powered by Fez SoftwareContact us