<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Wikipedia on English AI Terms Dictionary</title><link>https://terms-en.ai-term-hub.com/en/tags/wikipedia/</link><description>Recent content in Wikipedia on English AI Terms Dictionary</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sat, 18 Jul 2026 11:44:44 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/en/tags/wikipedia/index.xml" rel="self" type="application/rss+xml"/><item><title>Dataset:Natural Questions</title><link>https://terms-en.ai-term-hub.com/en/terms/datasetnatural_questions/</link><pubDate>Sat, 18 Jul 2026 09:53:44 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/en/terms/datasetnatural_questions/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Natural Questions (NQ) is a benchmark dataset introduced by Google to advance research in open-domain question answering. It maps real, anonymized search queries from Google to long-form answers found within Wikipedia articles. The dataset includes both &amp;lsquo;short answers&amp;rsquo; (specific spans of text) and &amp;rsquo;long answers&amp;rsquo; (paragraphs containing the short answer). It is essential for training models that can retrieve and synthesize information from vast knowledge bases to answer complex, real-world questions.&lt;/p></description></item><item><title>Dataset:Embedding Data/Simple Wiki</title><link>https://terms-en.ai-term-hub.com/en/terms/datasetembedding_datasimple_wiki/</link><pubDate>Sat, 18 Jul 2026 09:53:29 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/en/terms/datasetembedding_datasimple_wiki/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>This dataset consists of sentences and paragraphs extracted from Simple English Wikipedia, a version of Wikipedia written for non-native speakers with simplified grammar and vocabulary. It serves as a high-quality resource for training semantic embedding models, particularly those requiring robust generalization across diverse topics while maintaining linguistic simplicity. Researchers utilize it to benchmark how well models capture meaning in straightforward textual contexts, often improving performance on downstream tasks like classification and clustering where clarity is paramount.&lt;/p></description></item></channel></rss>