<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Data Engineering on English AI Terms Dictionary</title><link>https://terms-en.ai-term-hub.com/en/tags/data-engineering/</link><description>Recent content in Data Engineering on English AI Terms Dictionary</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sat, 18 Jul 2026 11:44:44 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/en/tags/data-engineering/index.xml" rel="self" type="application/rss+xml"/><item><title>Streaming</title><link>https://terms-en.ai-term-hub.com/en/terms/streaming/</link><pubDate>Sat, 18 Jul 2026 10:16:56 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/en/terms/streaming/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Streaming refers to the continuous ingestion and processing of data in real-time or near-real-time as it is generated. Unlike batch processing, which handles fixed datasets, streaming systems manage unbounded data flows with limited memory constraints. This requires algorithms capable of incremental updates and approximate results. Common technologies include Apache Kafka and Flink. Streaming is critical for applications requiring immediate insights, such as fraud detection, live monitoring, and dynamic recommendation engines, ensuring low latency and high throughput in distributed environments.&lt;/p></description></item><item><title>LlamaIndex</title><link>https://terms-en.ai-term-hub.com/en/terms/llamaindex/</link><pubDate>Sat, 18 Jul 2026 10:05:29 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/en/terms/llamaindex/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Originally known as GPT Index, LlamaIndex is a powerful data framework that enables LLMs to ingest and interact with structured and unstructured data. It provides tools for indexing, querying, and managing data pipelines, making it easier to build applications that leverage private or domain-specific information. By integrating seamlessly with various vector databases and embedding models, LlamaIndex simplifies the implementation of Retrieval-Augmented Generation (RAG), allowing developers to create context-aware AI assistants with minimal boilerplate code.&lt;/p></description></item><item><title>Knowledge integration</title><link>https://terms-en.ai-term-hub.com/en/terms/knowledge_integration/</link><pubDate>Sat, 18 Jul 2026 10:03:55 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/en/terms/knowledge_integration/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Knowledge integration involves merging data from diverse origins, such as databases, ontologies, and unstructured text, into a coherent schema. It addresses issues of semantic heterogeneity and inconsistency to create a single source of truth. This unified view enables more robust inference and decision-making by leveraging complementary information across different domains and formats.&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>The process of combining heterogeneous knowledge sources into a unified, consistent representation for enhanced reasoning.&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>Data fusion&lt;/li>
&lt;li>Ontology alignment&lt;/li>
&lt;li>Semantic interoperability&lt;/li>
&lt;li>Schema mapping&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>Enterprise data warehousing&lt;/li>
&lt;li>Multi-source medical diagnosis systems&lt;/li>
&lt;li>Integrating IoT sensor data with historical records&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/data-fusion/">Data Fusion&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/knowledge-graph/">Knowledge Graph&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/semantic-web/">Semantic Web&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/information-retrieval/">Information Retrieval&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Batch Processing</title><link>https://terms-en.ai-term-hub.com/en/terms/batch_processing/</link><pubDate>Sat, 18 Jul 2026 09:47:51 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/en/terms/batch_processing/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Batch processing involves aggregating data inputs into a group, or batch, before executing a computation or model inference. This approach contrasts with real-time streaming processing by allowing for higher throughput and better resource utilization through parallel execution. It is commonly used in offline training scenarios, historical data analysis, and scheduled tasks where immediate results are not required, optimizing hardware usage by maximizing GPU/TPU occupancy.&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>A computational method where data is collected over time and processed in groups rather than individually.&lt;/p></description></item></channel></rss>