diff --git a/src/data/roadmaps/data-engineer/content/amazon-ec2--compute@AHLsBfPfBJOhLlJ-64GcK.md b/src/data/roadmaps/data-engineer/content/amazon-ec2--compute@AHLsBfPfBJOhLlJ-64GcK.md index 249c4fdef..16f62978d 100644 --- a/src/data/roadmaps/data-engineer/content/amazon-ec2--compute@AHLsBfPfBJOhLlJ-64GcK.md +++ b/src/data/roadmaps/data-engineer/content/amazon-ec2--compute@AHLsBfPfBJOhLlJ-64GcK.md @@ -1,6 +1,6 @@ -# Amazon EC2 ( Compute) - -Amazon Elastic Compute Cloud (EC2) is a web service that provides secure, resizable compute capacity in the cloud. It is designed to make web-scale cloud computing easier for developers. EC2’s simple web service interface allows you to obtain and configure capacity with minimal friction. EC2 enables you to scale your compute capacity, develop and deploy applications faster, and run applications on AWS's reliable computing environment. You have the control of your computing resources and can access various configurations of CPU, Memory, Storage, and Networking capacity for your instances. +# Amazon EC2 (Compute) + +Amazon EC2 (Elastic Compute Cloud) provides virtual servers in the AWS cloud. Users can choose instance types optimized for compute, memory, or storage, and pay only for what they run. EC2 is used for running data processing jobs, hosting databases, and building custom data infrastructure on AWS. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/amazon-rds-database@GtFk7phYGfXUhxanicYNQ.md b/src/data/roadmaps/data-engineer/content/amazon-rds-database@GtFk7phYGfXUhxanicYNQ.md index 1f5789b53..17aa6d762 100644 --- a/src/data/roadmaps/data-engineer/content/amazon-rds-database@GtFk7phYGfXUhxanicYNQ.md +++ b/src/data/roadmaps/data-engineer/content/amazon-rds-database@GtFk7phYGfXUhxanicYNQ.md @@ -1,6 +1,6 @@ # Amazon RDS (Database) - -Amazon RDS (Relational Database Service) is a web service from Amazon Web Services. It's designed to simplify the setup, operation, and scaling of relational databases in the cloud. This service provides cost-efficient, resizable capacity for an industry-standard relational database and manages common database administration tasks. RDS supports six database engines: Amazon Aurora, PostgreSQL, MySQL, MariaDB, Oracle Database, and SQL Server. These engines give you the ability to run instances ranging from 5GB to 6TB of memory, accommodating your specific use case. It also ensures the database is up-to-date with the latest patches, automatically backs up your data and offers encryption at rest and in transit. + +Amazon RDS (Relational Database Service) is a managed relational database service from AWS that supports MySQL, PostgreSQL, MariaDB, Oracle, and MS SQL Server. It handles provisioning, backups, patching, and replication automatically. RDS is used for transactional databases that require minimal database administration overhead. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/apache-kafka@fTpx6m8U0506ZLCdDU5OG.md b/src/data/roadmaps/data-engineer/content/apache-kafka@fTpx6m8U0506ZLCdDU5OG.md index d44680ab3..9384a1f25 100644 --- a/src/data/roadmaps/data-engineer/content/apache-kafka@fTpx6m8U0506ZLCdDU5OG.md +++ b/src/data/roadmaps/data-engineer/content/apache-kafka@fTpx6m8U0506ZLCdDU5OG.md @@ -4,9 +4,7 @@ Apache Kafka is an open-source stream-processing software platform developed by Visit the following resources to learn more: -- [@official@Apache Kafka](https://kafka.apache.org/quickstart) -- [@article@Apache Kafka Streams](https://docs.confluent.io/platform/current/streams/concepts.html) +- [@official@Apache Kafka Docs](https://kafka.apache.org/43/getting-started/introduction/) - [@article@Kafka Streams Confluent](https://kafka.apache.org/documentation/streams/) - [@video@Apache Kafka Fundamentals](https://www.youtube.com/watch?v=B5j3uNBH8X4) -- [@video@Kafka in 100 Seconds](https://www.youtube.com/watch?v=uvb00oaa3k8) -- [@feed@Explore top posts about Kafka](https://app.daily.dev/tags/kafka?ref=roadmapsh) \ No newline at end of file +- [@video@Kafka in 100 Seconds](https://www.youtube.com/watch?v=uvb00oaa3k8) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/apache-spark@qHMtJFYcGmESiz_VwRwiI.md b/src/data/roadmaps/data-engineer/content/apache-spark@qHMtJFYcGmESiz_VwRwiI.md index 8f6619d07..ead9d2fdc 100644 --- a/src/data/roadmaps/data-engineer/content/apache-spark@qHMtJFYcGmESiz_VwRwiI.md +++ b/src/data/roadmaps/data-engineer/content/apache-spark@qHMtJFYcGmESiz_VwRwiI.md @@ -1,9 +1,8 @@ # Apache Spark - -Apache Spark is an open-source distributed computing system designed for big data processing and analytics. It offers a unified interface for programming entire clusters, enabling efficient handling of large-scale data with built-in support for data parallelism and fault tolerance. Spark excels in processing tasks like batch processing, real-time data streaming, machine learning, and graph processing. It’s known for its speed, ease of use, and ability to process data in-memory, significantly outperforming traditional MapReduce systems. Spark is widely used in big data ecosystems for its scalability and versatility across various data processing tasks. + +Apache Spark is a distributed data processing engine for large-scale batch and streaming workloads. It processes data in memory across a cluster, making it significantly faster than MapReduce for many workloads. Spark supports Python, Scala, Java, and R, and provides APIs for SQL, streaming, machine learning, and graph processing. Visit the following resources to learn more: - [@official@ApacheSpark](https://spark.apache.org/documentation.html) -- [@article@Spark By Examples](https://sparkbyexamples.com) -- [@feed@Explore top posts about Apache Spark](https://app.daily.dev/tags/spark?ref=roadmapsh) \ No newline at end of file +- [@article@Spark By Examples](https://sparkbyexamples.com) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/apis@cxTriSZvrmXP4axKynIZW.md b/src/data/roadmaps/data-engineer/content/apis@cxTriSZvrmXP4axKynIZW.md index 1e8d62755..e0a60a8a4 100644 --- a/src/data/roadmaps/data-engineer/content/apis@cxTriSZvrmXP4axKynIZW.md +++ b/src/data/roadmaps/data-engineer/content/apis@cxTriSZvrmXP4axKynIZW.md @@ -1,8 +1,9 @@ -# APIs and Data Collection - -Application Programming Interfaces, better known as APIs, play a fundamental role in the work of data engineers, particularly in the process of data collection. APIs are sets of protocols, routines, and tools that enable different software applications to communicate with each other. An API allows developers to interact with a service or platform through a defined set of rules and endpoints, enabling data exchange and functionality use without needing to understand the underlying code. In data engineering, APIs are used extensively to collect, exchange, and manipulate data from different sources in a secure and efficient manner. +# APIs + +APIs (Application Programming Interfaces) expose data from external services in a structured format, typically JSON or XML over HTTP. Many data pipelines pull data from third-party APIs such as payment processors, marketing platforms, or social networks. Rate limits, authentication, and schema changes are common challenges when ingesting from APIs. Visit the following resources to learn more: +- [@roadmap@Visit the Dedicated Java Roadmap](https://roadmap.sh/api-design) - [@article@What is an API?](https://aws.amazon.com/what-is/api/) - [@article@A Beginner's Guide to APIs](https://www.postman.com/what-is-an-api/) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/argocd@PUzHbjwntTSj1REL_dAov.md b/src/data/roadmaps/data-engineer/content/argocd@PUzHbjwntTSj1REL_dAov.md index e2310a4bc..aec5b8b0b 100644 --- a/src/data/roadmaps/data-engineer/content/argocd@PUzHbjwntTSj1REL_dAov.md +++ b/src/data/roadmaps/data-engineer/content/argocd@PUzHbjwntTSj1REL_dAov.md @@ -1,10 +1,9 @@ # ArgoCD - -Argo CD is a continuous delivery tool for Kubernetes that is based on the GitOps methodology. It is used to automate the deployment and management of cloud-native applications by continuously synchronizing the desired application state with the actual application state in the production environment. In an Argo CD workflow, changes to the application are made by committing code or configuration changes to a Git repository. Argo CD monitors the repository and automatically deploys the changes to the production environment using a continuous delivery pipeline. The pipeline is triggered by changes to the Git repository and is responsible for building, testing, and deploying the changes to the production environment. Argo CD is designed to be a simple and efficient way to manage cloud-native applications, as it allows developers to make changes to the system using familiar tools and processes and it provides a clear and auditable history of all changes to the system. It is often used in conjunction with tools such as Helm to automate the deployment and management of cloud-native applications. + +Argo CD is a declarative GitOps continuous delivery tool for Kubernetes. It continuously monitors a Git repository and ensures that the state of the Kubernetes cluster matches the desired state defined in code. Argo CD is used in data engineering to manage Kubernetes-based pipeline deployments and infrastructure changes through Git. Visit the following resources to learn more: - [@official@Argo CD - Argo Project](https://argo-cd.readthedocs.io/en/stable/) - [@video@ArgoCD Tutorial for Beginners](https://www.youtube.com/watch?v=MeU5_k9ssrs) -- [@video@What is ArgoCD](https://www.youtube.com/watch?v=p-kAqxuJNik) -- [@feed@Explore top posts about ArgoCD](https://app.daily.dev/tags/argocd?ref=roadmapsh) \ No newline at end of file +- [@video@What is ArgoCD](https://www.youtube.com/watch?v=p-kAqxuJNik) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/async-vs-sync-communication@VefHaP7rIOcZVFzglyn66.md b/src/data/roadmaps/data-engineer/content/async-vs-sync-communication@VefHaP7rIOcZVFzglyn66.md index 69d7ed421..8e8b2bbc6 100644 --- a/src/data/roadmaps/data-engineer/content/async-vs-sync-communication@VefHaP7rIOcZVFzglyn66.md +++ b/src/data/roadmaps/data-engineer/content/async-vs-sync-communication@VefHaP7rIOcZVFzglyn66.md @@ -1,8 +1,6 @@ # Async vs Sync Communication - -Synchronous and asynchronous data refer to different approaches in data transmission and processing. **Synchronous** ingestion is a process where the system waits for a response from the data source before proceeding. In contrast, **asynchronous** ingestion is a process where data is ingested without waiting for a response from the data source. Normally, data is queued in a buffer and sent in batches for efficiency. - -Each approach has its benefits and drawbacks, and the choice depends on the specific requirements of the data ingestion process and the business needs. + +Synchronous communication means the sender waits for a response before continuing. Asynchronous communication means the sender sends a message and continues without waiting. Messaging systems enable asynchronous communication, which is better suited for high-throughput pipelines where blocking would create bottlenecks. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/aurora-db@YZ4G1-6VJ7VdsphdcBTf9.md b/src/data/roadmaps/data-engineer/content/aurora-db@YZ4G1-6VJ7VdsphdcBTf9.md index 330ff8357..685129e0f 100644 --- a/src/data/roadmaps/data-engineer/content/aurora-db@YZ4G1-6VJ7VdsphdcBTf9.md +++ b/src/data/roadmaps/data-engineer/content/aurora-db@YZ4G1-6VJ7VdsphdcBTf9.md @@ -4,5 +4,4 @@ Amazon Aurora (Aurora) is a fully managed relational database engine that's comp Visit the following resources to learn more: -- [@official@SAmazon Aurora](https://aws.amazon.com/rds/aurora/) -- [@article@SAmazon Aurora: What It Is, How It Works, and How to Get Started](https://www.datacamp.com/tutorial/amazon-aurora) \ No newline at end of file +- [@official@SAmazon Aurora](https://aws.amazon.com/rds/aurora/) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/aws-eks@eVqcYI2Sy2Dldl3SfxB2C.md b/src/data/roadmaps/data-engineer/content/aws-eks@eVqcYI2Sy2Dldl3SfxB2C.md index 57df891f7..883ef00c4 100644 --- a/src/data/roadmaps/data-engineer/content/aws-eks@eVqcYI2Sy2Dldl3SfxB2C.md +++ b/src/data/roadmaps/data-engineer/content/aws-eks@eVqcYI2Sy2Dldl3SfxB2C.md @@ -1,6 +1,6 @@ -# EKS - -Amazon Elastic Kubernetes Service (EKS) is a managed service that simplifies the deployment, management, and scaling of containerized applications using Kubernetes, an open-source container orchestration platform. EKS manages the Kubernetes control plane for the user, making it easy to run Kubernetes applications without the operational overhead of maintaining the Kubernetes control plane. With EKS, you can leverage AWS services such as Auto Scaling Groups, Elastic Load Balancer, and Route 53 for resilient and scalable application infrastructure. Additionally, EKS can support Spot and On-Demand instances use, and includes integrations with AWS App Mesh service and AWS Fargate for serverless compute. +# AWS EKS + +Amazon EKS (Elastic Kubernetes Service) is AWS's managed Kubernetes service. It runs the Kubernetes control plane across multiple availability zones and integrates with AWS services like IAM, VPC, and ECR. EKS is used to run containerized data pipelines and services on AWS without managing the Kubernetes control plane directly. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/aws-sns@uFeiTRobSymkvCinhwmZV.md b/src/data/roadmaps/data-engineer/content/aws-sns@uFeiTRobSymkvCinhwmZV.md index 84c009ed7..c3b7789b3 100644 --- a/src/data/roadmaps/data-engineer/content/aws-sns@uFeiTRobSymkvCinhwmZV.md +++ b/src/data/roadmaps/data-engineer/content/aws-sns@uFeiTRobSymkvCinhwmZV.md @@ -1,9 +1,9 @@ # AWS SNS - -Amazon Simple Notification Service (Amazon SNS) is a web service that makes it easy to set up, operate, and send notifications from the cloud. It provides developers with a highly scalable, flexible, and cost-effective capability to publish messages from an application and immediately deliver them to subscribers or other applications. It is designed to make web-scale computing easier for developers. Amazon SNS follows the “publish-subscribe” (pub-sub) messaging paradigm, with notifications being delivered to clients using a “push” mechanism that eliminates the need to periodically check or “poll” for new information and updates. With simple APIs requiring minimal up-front development effort, no maintenance or management overhead and pay-as-you-go pricing, Amazon SNS gives developers an easy mechanism to incorporate a powerful notification system with their applications. + +Amazon SNS (Simple Notification Service) is a fully managed pub/sub messaging service from AWS. It allows a single message to be sent to multiple subscribers simultaneously through topics. SNS is commonly used alongside SQS to fan out messages to multiple queues or trigger downstream processing in Lambda functions and data pipelines. Visit the following resources to learn more: -- [@official@Amazon Simple Notification Service (SNS) ](http://aws.amazon.com/sns/) +- [@official@Amazon Simple Notification Service (SNS)](http://aws.amazon.com/sns/) - [@official@Send Fanout Event Notifications](https://aws.amazon.com/getting-started/hands-on/send-fanout-event-notifications/) - [@article@What is Pub/Sub Messaging?](https://aws.amazon.com/what-is/pub-sub-messaging/) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/aws-sqs@uIU5Yncp6hGDcNO1fpjUS.md b/src/data/roadmaps/data-engineer/content/aws-sqs@uIU5Yncp6hGDcNO1fpjUS.md index da641452c..9bb2c425b 100644 --- a/src/data/roadmaps/data-engineer/content/aws-sqs@uIU5Yncp6hGDcNO1fpjUS.md +++ b/src/data/roadmaps/data-engineer/content/aws-sqs@uIU5Yncp6hGDcNO1fpjUS.md @@ -5,5 +5,4 @@ Amazon Simple Queue Service (Amazon SQS) offers a secure, durable, and available Visit the following resources to learn more: - [@official@Amazon Simple Queue Service](https://aws.amazon.com/sqs/) -- [@official@What is Amazon Simple Queue Service?](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/welcome.html) -- [@article@Amazon Simple Queue Service (SQS): A Comprehensive Tutorial](https://www.datacamp.com/tutorial/amazon-sqs) \ No newline at end of file +- [@official@What is Amazon Simple Queue Service?](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/welcome.html) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/azure-sql-database@iIZ3g70KRwEJCBNaONd2d.md b/src/data/roadmaps/data-engineer/content/azure-sql-database@iIZ3g70KRwEJCBNaONd2d.md index 26c9901ca..5e6ba96e0 100644 --- a/src/data/roadmaps/data-engineer/content/azure-sql-database@iIZ3g70KRwEJCBNaONd2d.md +++ b/src/data/roadmaps/data-engineer/content/azure-sql-database@iIZ3g70KRwEJCBNaONd2d.md @@ -6,5 +6,4 @@ Visit the following resources to learn more: - [@official@Azure SQL Database](https://azure.microsoft.com/en-us/products/azure-sql/database) - [@official@What is Azure SQL Database?](https://learn.microsoft.com/en-us/azure/azure-sql/database/sql-database-paas-overview?view=azuresql) -- [@article@Azure SQL Database: Step-by-Step Setup and Management](https://www.datacamp.com/tutorial/azure-sql-database) - [@video@Azure SQL for Beginners](https://www.youtube.com/playlist?list=PLlrxD0HtieHi5c9-i_Dnxw9vxBY-TqaeN) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/best-practices@yyJJGinOv3M21MFuqJs0j.md b/src/data/roadmaps/data-engineer/content/best-practices@yyJJGinOv3M21MFuqJs0j.md index 1ccac39ca..a0fba7589 100644 --- a/src/data/roadmaps/data-engineer/content/best-practices@yyJJGinOv3M21MFuqJs0j.md +++ b/src/data/roadmaps/data-engineer/content/best-practices@yyJJGinOv3M21MFuqJs0j.md @@ -1,14 +1,6 @@ # Best Practices - -1. **Ensure Reliability.** A robust messaging system must guarantee that messages aren’t lost, even during node failures or network issues. This means using acknowledgments, replication across multiple brokers, and durable storage on disk. These measures ensure that producers and consumers can recover seamlessly without data loss when something goes wrong. - -2. **Design for Scalability.** Scalability should be baked in from the start. Partition topics strategically to distribute load across brokers and consumer groups, enabling horizontal scaling. - -3. **Maintain Message Ordering.** For systems that depend on message sequence, ensure ordering within partitions and design producers to consistently route related messages to the same partition. - -4. **Secure Communication.** Messaging queues often carry sensitive data, so encrypt messages both in transit and at rest. Implement authentication techniques to ensure only trusted clients can publish or consume, and enforce authorization rules to limit access to specific topics or operations. - -5. **Monitor & Alert.** Continuous visibility into your messaging system is essential. Track metrics such as message lag, throughput, consumer group health, and broker disk usage. Set alerts for abnormal patterns, like growing lag or dropped connections, so you can respond before they affect downstream systems. + +Best practices for messaging systems include designing idempotent consumers to handle duplicate delivery, setting appropriate retention and replication policies, monitoring consumer lag, and planning for schema evolution. Dead-letter queues are used to handle messages that fail processing repeatedly without losing them. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/big-data-tools@03BHmPhYkZrJwRvQdmxxr.md b/src/data/roadmaps/data-engineer/content/big-data-tools@03BHmPhYkZrJwRvQdmxxr.md index 002424033..85910d27a 100644 --- a/src/data/roadmaps/data-engineer/content/big-data-tools@03BHmPhYkZrJwRvQdmxxr.md +++ b/src/data/roadmaps/data-engineer/content/big-data-tools@03BHmPhYkZrJwRvQdmxxr.md @@ -1,11 +1,8 @@ # Big Data Tools - -Big data tools are specialized software and platforms designed to handle the massive volume, velocity, and variety of data that traditional data processing tools cannot effectively manage. These tools provide the infrastructure, frameworks, and capabilities to process, analyze, and extract meaningful knowledge from vast datasets. They are essential for modern data-driven organizations seeking to gain insights, make informed decisions, and achieve a competitive advantage. - -Hadoop and Spark are two of the most prominent frameworks in big data they handle the processing of large-scale data in very different ways. While Hadoop can be credited with democratizing the distributed computing paradigm through a robust storage system called HDFS and a computational model called MapReduce, Spark is changing the game with its in-memory architecture and flexible programming model. + +Big data tools are designed to process and analyze datasets too large to handle with traditional single-machine tools. They distribute computation across clusters and are optimized for throughput at scale. The most widely used big data processing framework is Apache Spark, with Hadoop-based tools remaining common in legacy environments. Visit the following resources to learn more: - [@article@What is Big Data?](https://cloud.google.com/learn/what-is-big-data?hl=en) -- [@article@Hadoop vs Spark: Which Big Data Framework Is Right For You?](https://www.datacamp.com/blog/hadoop-vs-spark) -- [@video@introduction to Big Data with Spark and Hadoop](http://youtube.com/watch?v=vHlwg4ciCsI&t=80s&ab_channel=freeCodeAcademy) \ No newline at end of file +- [@video@Introduction to Big Data with Spark and Hadoop](http://youtube.com/watch?v=vHlwg4ciCsI&t=80s&ab_channel=freeCodeAcademy) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/bigtable@ltZftFsiOo12AkQ-04N3B.md b/src/data/roadmaps/data-engineer/content/bigtable@ltZftFsiOo12AkQ-04N3B.md index 344cf5721..e046a610e 100644 --- a/src/data/roadmaps/data-engineer/content/bigtable@ltZftFsiOo12AkQ-04N3B.md +++ b/src/data/roadmaps/data-engineer/content/bigtable@ltZftFsiOo12AkQ-04N3B.md @@ -1,6 +1,6 @@ # BigTable -Bigtable is a high-performance, scalable database that excels at capturing, processing, and analyzing data in real-time. It aggregates data as it's written, providing immediate insights into user behavior, A/B testing results, and engagement metrics. This real-time capability also fuels AI/ML models for interactive applications. Bigtable integrates seamlessly with both Dataflow, enriching streaming pipelines with low-latency lookups, and BigQuery, enabling real-time serving of analytics in user facing application and ad-hoc querying on the same data. +Bigtable is a high-performance, scalable database that excels at capturing, processing, and analyzing data in real-time. It aggregates data as it's written, providing immediate insights into user behavior, A/B testing results, and engagement metrics. This real-time capability also fuels AI/ML models for interactive applications. Bigtable integrates seamlessly with both Dataflow, enriching streaming pipelines with low-latency lookups, and BigQuery, enabling real-time serving of analytics in user-facing applications and ad-hoc querying on the same data. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/business-intelligence@zA5QqqBMsqymdiPGFdUnt.md b/src/data/roadmaps/data-engineer/content/business-intelligence@zA5QqqBMsqymdiPGFdUnt.md index 66ee212f3..f4cab7a5e 100644 --- a/src/data/roadmaps/data-engineer/content/business-intelligence@zA5QqqBMsqymdiPGFdUnt.md +++ b/src/data/roadmaps/data-engineer/content/business-intelligence@zA5QqqBMsqymdiPGFdUnt.md @@ -1,11 +1,10 @@ # Business Intelligence - -Business intelligence encompasses a set of techniques and technologies to transform raw data into meaningful insights that drive strategic decision-making within an organization. BI tools enable business users to access different types of data, historical and current, third-party and in-house, as well as semistructured data and unstructured data such as social media. Users can analyze this information to gain insights into how the business is performing and what it should do next. - -BI platforms traditionally rely on data warehouses for their baseline information. The strength of a data warehouse is that it aggregates data from multiple data sources into one central system to support business data analytics and reporting. BI presents the results to the user in the form of reports, charts and maps, which might be displayed through a dashboard. + +Business intelligence (BI) refers to the tools and processes used to collect, analyze, and visualize business data to support decisions. BI platforms connect to data warehouses and allow business users to build reports and dashboards without writing code. Common BI tools include Tableau, Power BI, Looker, and Streamlit. Visit the following resources to learn more: +- [@roadmap@Visit the Dedicated BI Analyst Roadmap](https://roadmap.sh/bi-analyst) - [@article@What is business intelligence (BI)?](https://www.ibm.com/think/topics/business-intelligence) - [@article@Business intelligence: A complete overview](https://www.tableau.com/business-intelligence/what-is-business-intelligence) - [@video@What is business intelligence?](https://www.youtube.com/watch?v=l98-BcB3UIE) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/cap-theorem@AslPFjoakcC44CmPB5nuw.md b/src/data/roadmaps/data-engineer/content/cap-theorem@AslPFjoakcC44CmPB5nuw.md index e8d19f8da..f038edbf1 100644 --- a/src/data/roadmaps/data-engineer/content/cap-theorem@AslPFjoakcC44CmPB5nuw.md +++ b/src/data/roadmaps/data-engineer/content/cap-theorem@AslPFjoakcC44CmPB5nuw.md @@ -1,6 +1,6 @@ # CAP Theorem - -The CAP Theorem, also known as Brewer's Theorem, is a fundamental principle in distributed database systems. It states that in a distributed system, it's impossible to simultaneously guarantee all three of the following properties: Consistency (all nodes see the same data at the same time), Availability (every request receives a response, without guarantee that it contains the most recent version of the data), and Partition tolerance (the system continues to operate despite network failures between nodes). According to the theorem, a distributed system can only strongly provide two of these three guarantees at any given time. This principle guides the design and architecture of distributed systems, influencing decisions on data consistency models, replication strategies, and failure handling. Understanding the CAP Theorem is crucial for designing robust, scalable distributed systems and for choosing appropriate database solutions for specific use cases in distributed computing environments. + +The CAP theorem states that a distributed system can provide at most two of three guarantees: Consistency, Availability, and Partition Tolerance. In practice, network partitions are unavoidable, so systems must choose between consistency and availability when a partition occurs. This trade-off shapes the design of distributed databases like Cassandra, DynamoDB, and HBase. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/cassandra@QYR8ESN7xhi4ZxcoiZbgn.md b/src/data/roadmaps/data-engineer/content/cassandra@QYR8ESN7xhi4ZxcoiZbgn.md index 92291bb39..c000d8ada 100644 --- a/src/data/roadmaps/data-engineer/content/cassandra@QYR8ESN7xhi4ZxcoiZbgn.md +++ b/src/data/roadmaps/data-engineer/content/cassandra@QYR8ESN7xhi4ZxcoiZbgn.md @@ -1,10 +1,9 @@ # Cassandra - -Apache Cassandra is a highly scalable, distributed NoSQL database designed to handle large amounts of structured data across multiple commodity servers. It provides high availability with no single point of failure, offering linear scalability and proven fault-tolerance on commodity hardware or cloud infrastructure. Cassandra uses a masterless ring architecture, where all nodes are equal, allowing for easy data distribution and replication. It supports flexible data models and can handle both unstructured and structured data. Cassandra excels in write-heavy environments and is particularly suitable for applications requiring high throughput and low latency. Its data model is based on wide column stores, offering a more complex structure than key-value stores. Widely used in big data applications, Cassandra is known for its ability to handle massive datasets while maintaining performance and reliability. + +Apache Cassandra is an open-source distributed wide-column database designed for high availability and linear scalability. It has no single point of failure and is optimized for fast writes across multiple data centers. Cassandra is used for time-series data, IoT workloads, and applications requiring continuous uptime. Visit the following resources to learn more: - [@official@Apache Cassandra](https://cassandra.apache.org/_/index.html) -- [@article@article@Cassandra - Quick Guide](https://www.tutorialspoint.com/cassandra/cassandra_quick_guide.htm) -- [@video@Apache Cassandra - Course for Beginners](https://www.youtube.com/watch?v=J-cSy5MeMOA) -- [@feed@Explore top posts about Backend Development](https://app.daily.dev/tags/backend?ref=roadmapsh) \ No newline at end of file +- [@article@Cassandra - Quick Guide](https://www.tutorialspoint.com/cassandra/cassandra_quick_guide.htm) +- [@video@Apache Cassandra - Course for Beginners](https://www.youtube.com/watch?v=J-cSy5MeMOA) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/census@vZGDtlyt_yj4szcPTw3cv.md b/src/data/roadmaps/data-engineer/content/census@vZGDtlyt_yj4szcPTw3cv.md index 3b5748f3e..31f0d536c 100644 --- a/src/data/roadmaps/data-engineer/content/census@vZGDtlyt_yj4szcPTw3cv.md +++ b/src/data/roadmaps/data-engineer/content/census@vZGDtlyt_yj4szcPTw3cv.md @@ -4,7 +4,6 @@ Census is a reverse ETL platform that synchronizes data from a data warehouse to Visit the following resources to learn more: -- [@official@Census](https://www.getcensus.com/reverse-etl) - [@official@Census Documentation](https://developers.getcensus.com/getting-started/introduction) - [@article@A starter guide to reverse ETL with Census](https://www.getcensus.com/blog/starter-guide-for-first-time-census-users) - [@video@How to "Reverse ETL" with Census](https://www.youtube.com/watch?v=XkS7DQFHzbA) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/choosing-the-right-technologies@_MpdVlvvkrsgzigYMZ_P8.md b/src/data/roadmaps/data-engineer/content/choosing-the-right-technologies@_MpdVlvvkrsgzigYMZ_P8.md index de263b7ce..a1b5583da 100644 --- a/src/data/roadmaps/data-engineer/content/choosing-the-right-technologies@_MpdVlvvkrsgzigYMZ_P8.md +++ b/src/data/roadmaps/data-engineer/content/choosing-the-right-technologies@_MpdVlvvkrsgzigYMZ_P8.md @@ -1,13 +1,6 @@ # Choosing the Right Technologies - -The data engineering ecosystem is rapidly expanding, and selecting the right technologies for your use case can be challenging. Below you can find some considerations for choosing data technologies across the data engineering lifecycle: - -* **Team size and capabilities.** Your team's size will determine the amount of bandwidth your team can dedicate to complex solutions. For small teams, try to stick to simple solutions and technologies your team is familiar with. -* **Interoperability**. When choosing a technology or system, you’ll need to ensure that it interacts and operates smoothly with other technologies. -* **Cost optimization and business value,** Consider direct and indirect costs of a technology and the opportunity cost of choosing some technologies over others. -* **Location** Companies have many options when it comes to choosing where to run their technology stack, including cloud providers, on-premises systems, hybrid clouds, and multicloud. -* **Build versus buy**. Depending on your needs and capabilities, you can either invest in building your own technologies, implement open-source solutions, or purchase proprietary solutions and services. -* **Server versus serverless**. Depending on your needs, you may prefer server-based setups, where developers manage servers, or serverless systems, which translates the server management to cloud providers, allowing developers to focus solely on writing code. + +Selecting the right technology stack depends on data volume, team size, latency requirements, and budget. There is no universal best choice; a small startup may do well with a simple Postgres setup, while a large enterprise may need distributed processing and a cloud data warehouse. The decision involves evaluating trade-offs between cost, complexity, scalability, and maintainability. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/cicd@k2SJ4ELGa4B2ZERDAk1uj.md b/src/data/roadmaps/data-engineer/content/cicd@k2SJ4ELGa4B2ZERDAk1uj.md index bb39d802e..a4485e7f6 100644 --- a/src/data/roadmaps/data-engineer/content/cicd@k2SJ4ELGa4B2ZERDAk1uj.md +++ b/src/data/roadmaps/data-engineer/content/cicd@k2SJ4ELGa4B2ZERDAk1uj.md @@ -1,8 +1,6 @@ -# CI / CD - -**Continuous Integration** is a software development method where team members integrate their work at least once daily. An automated build checks every integration to detect errors in this method. In Continuous Integration, the software is built and tested immediately after a code commit. In a large project with many developers, commits are made many times during the day. With each commit, code is built and tested. - -**Continuous Delivery** is a software engineering method in which a team develops software products in a short cycle. It ensures that software can be easily released at any time. The main aim of continuous delivery is to build, test, and release software with good speed and frequency. It helps reduce the cost, time, and risk of delivering changes by allowing for frequent updates in production. +# CI/CD + +CI/CD (Continuous Integration and Continuous Delivery) is a set of practices and tools for automating the testing and deployment of code changes. In data engineering, CI/CD pipelines validate pipeline code, run tests, and deploy updates to production automatically. This reduces manual errors and accelerates the delivery of pipeline changes. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/circle-ci@CewITBPtfVs32LD5Acb2E.md b/src/data/roadmaps/data-engineer/content/circle-ci@CewITBPtfVs32LD5Acb2E.md index e0a4260b0..779eae789 100644 --- a/src/data/roadmaps/data-engineer/content/circle-ci@CewITBPtfVs32LD5Acb2E.md +++ b/src/data/roadmaps/data-engineer/content/circle-ci@CewITBPtfVs32LD5Acb2E.md @@ -6,5 +6,4 @@ Visit the following resources to learn more: - [@official@CircleCI](https://circleci.com/) - [@official@CircleCI Documentation](https://circleci.com/docs) -- [@official@Configuration Tutorial](https://circleci.com/docs/config-intro) -- [@feed@Explore top posts about CI/CD](https://app.daily.dev/tags/cicd?ref=roadmapsh) \ No newline at end of file +- [@official@Configuration Tutorial](https://circleci.com/docs/config-intro) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/cloud-architectures@YLfyb_ycgz1hu0yW8SPNE.md b/src/data/roadmaps/data-engineer/content/cloud-architectures@YLfyb_ycgz1hu0yW8SPNE.md index 8ac5dac42..019617f65 100644 --- a/src/data/roadmaps/data-engineer/content/cloud-architectures@YLfyb_ycgz1hu0yW8SPNE.md +++ b/src/data/roadmaps/data-engineer/content/cloud-architectures@YLfyb_ycgz1hu0yW8SPNE.md @@ -1,13 +1,6 @@ # Cloud Architectures - -Cloud architecture refers to how various cloud technology components, such as hardware, virtual resources, software capabilities, and virtual network systems interact and connect to create cloud computing environments. Cloud architecture dictates how components are integrated so that you can pool, share, and scale resources over a network. It acts as a blueprint that defines the best way to strategically combine resources to build a cloud environment for a specific business need. - -Cloud architecture components can included, among others: - -* A frontend platform -* A backend platform -* A cloud-based delivery model -* A network (internet, intranet, or intercloud) + +Cloud architectures describe how systems are designed to run on cloud infrastructure. Common patterns include multi-tier architectures, microservices, event-driven designs, and serverless functions. Good cloud architecture balances cost, reliability, scalability, and security. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/cloud-computing@lDeSL9qvgQgyAMcWXF7Fr.md b/src/data/roadmaps/data-engineer/content/cloud-computing@lDeSL9qvgQgyAMcWXF7Fr.md index 3855f70fa..ee47ae068 100644 --- a/src/data/roadmaps/data-engineer/content/cloud-computing@lDeSL9qvgQgyAMcWXF7Fr.md +++ b/src/data/roadmaps/data-engineer/content/cloud-computing@lDeSL9qvgQgyAMcWXF7Fr.md @@ -1,6 +1,6 @@ # Cloud Computing - -**Cloud Computing** refers to the delivery of computing services over the internet rather than using local servers or personal devices. These services include servers, storage, databases, networking, software, analytics, and intelligence. Cloud Computing enables faster innovation, flexible resources, and economies of scale. There are various types of cloud computing such as public clouds, private clouds, and hybrids clouds. Furthermore, it's divided into different services like Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS). These services differ mainly in the level of control an organization has over their data and infrastructures. + +Cloud computing refers to the delivery of computing resources, including servers, storage, databases, networking, and software, over the internet. Major cloud providers offer on-demand infrastructure that scales with usage and is billed per consumption. Cloud platforms are the dominant environment for modern data engineering work. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/cluster-computing-basics@hB0y8A2U3owpAbTUb7LN5.md b/src/data/roadmaps/data-engineer/content/cluster-computing-basics@hB0y8A2U3owpAbTUb7LN5.md index ef8c25426..511a1c0ff 100644 --- a/src/data/roadmaps/data-engineer/content/cluster-computing-basics@hB0y8A2U3owpAbTUb7LN5.md +++ b/src/data/roadmaps/data-engineer/content/cluster-computing-basics@hB0y8A2U3owpAbTUb7LN5.md @@ -1,3 +1,3 @@ # Cluster Computing Basics - -Cluster computing is the process of using multiple computing nodes, called clusters, to increase processing power for solving complex problems, such as Big Data analytics and AI model training. These tasks require parallel processing of millions of data points for complex classification and prediction tasks. Cluster computing technology coordinates multiple computing nodes, each with its own CPUs, GPUs, and internal memory, to work together on the same data processing task. Applications on cluster computing infrastructure run as if on a single machine and are unaware of the underlying system complexities. \ No newline at end of file + +Cluster computing refers to using a group of connected machines that work together as a single system to process data. It enables workloads that are too large or slow for a single machine by distributing computation across multiple nodes. Concepts like job scheduling, distributed file systems, and resource management are central to working with clusters. \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/cluster-management-tools@wpZfbIFtfiUSLMASk4t7f.md b/src/data/roadmaps/data-engineer/content/cluster-management-tools@wpZfbIFtfiUSLMASk4t7f.md index 375b4a210..a528c84b9 100644 --- a/src/data/roadmaps/data-engineer/content/cluster-management-tools@wpZfbIFtfiUSLMASk4t7f.md +++ b/src/data/roadmaps/data-engineer/content/cluster-management-tools@wpZfbIFtfiUSLMASk4t7f.md @@ -1,5 +1,3 @@ # Cluster Management Tools -Cluster management software maximizes the work that a cluster of computers can perform. A cluster manager balances workload to reduce bottlenecks, monitors the health of the elements of the cluster, and manages failover when an element fails. A cluster manager can also help a system administrator to perform administration tasks on elements in the cluster. - -Some of the most popular Cluster Management Tools are Kubernetes and Apache Hadoop YARN. \ No newline at end of file +Cluster management software maximizes the work that a cluster of computers can perform. A cluster manager balances workload to reduce bottlenecks, monitors the health of the elements of the cluster, and manages failover when an element fails. A cluster manager can also help a system administrator to perform administration tasks on elements in the cluster. \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/column@fBD6ZQoMac8w4kMJw_Jrd.md b/src/data/roadmaps/data-engineer/content/column@fBD6ZQoMac8w4kMJw_Jrd.md index 8b55248fa..bba00b5d0 100644 --- a/src/data/roadmaps/data-engineer/content/column@fBD6ZQoMac8w4kMJw_Jrd.md +++ b/src/data/roadmaps/data-engineer/content/column@fBD6ZQoMac8w4kMJw_Jrd.md @@ -1,6 +1,6 @@ # Column - -A columnar database is a type of No-SQL database that stores data by columns instead of by rows. In a traditional SQL database, all the information for one record is stored together, but in a columnar database, all the values for a single column are stored together. This makes it much faster to read and analyze large amounts of data, especially when you only need a few columns instead of the whole record. For example, if you want to quickly find the average sales price from millions of rows, a columnar database can scan just the "price" column instead of every piece of data. This design is often used in data warehouses and analytics systems because it speeds up queries and saves storage space through better compression. + +Column-family databases (also called wide-column stores) organize data into rows and dynamic columns grouped into column families. They are optimized for read and write operations on large datasets spread across many machines. This model is well suited for time-series data, logging, and analytical workloads. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/compute-engine-compute@-cU86vJWJmlmPHXDCo31o.md b/src/data/roadmaps/data-engineer/content/compute-engine-compute@-cU86vJWJmlmPHXDCo31o.md index bf3f9ce45..f04886a3a 100644 --- a/src/data/roadmaps/data-engineer/content/compute-engine-compute@-cU86vJWJmlmPHXDCo31o.md +++ b/src/data/roadmaps/data-engineer/content/compute-engine-compute@-cU86vJWJmlmPHXDCo31o.md @@ -1,6 +1,6 @@ # Compute Engine (Compute) - -Compute Engine is a computing and hosting service that lets you create and run virtual machines on Google infrastructure. Compute Engine offers scale, performance, and value that lets you easily launch large compute clusters on Google's infrastructure. There are no upfront investments, and you can run thousands of virtual CPUs on a system that offers quick, consistent performance. You can configure and control Compute Engine resources using the Google Cloud console, the Google Cloud CLI, or using a REST-based API. You can also use a variety of programming languages to run Compute Engine, including Python, Go, and Java. + +Google Cloud Compute Engine provides virtual machine instances on Google's infrastructure. It supports custom machine types, preemptible VMs for cost savings, and integration with other Google Cloud services. Compute Engine is used for running custom workloads, data processing jobs, and services that require full control over the operating environment. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/containers--orchestration@eTHitN2erd6z8-MZiXE9s.md b/src/data/roadmaps/data-engineer/content/containers--orchestration@eTHitN2erd6z8-MZiXE9s.md index de04c61f1..7c551444b 100644 --- a/src/data/roadmaps/data-engineer/content/containers--orchestration@eTHitN2erd6z8-MZiXE9s.md +++ b/src/data/roadmaps/data-engineer/content/containers--orchestration@eTHitN2erd6z8-MZiXE9s.md @@ -1,14 +1,10 @@ # Containers & Orchestration - -**Containers** are lightweight, portable, and isolated environments that package applications and their dependencies, enabling consistent deployment across different computing environments. They encapsulate software code, runtime, system tools, libraries, and settings, ensuring that the application runs the same regardless of where it's deployed. Containers share the host operating system's kernel, making them more efficient than traditional virtual machines. - -**Orchestration** refers to the automated coordination and management of complex IT systems. It involves combining multiple automated tasks and processes into a single workflow to achieve a specific goal. Orchestration is one of the key components of any software development process and it should never be avoided nor preferred over manual configuration. As an automation practice, orchestration helps to remove the chance of human error from the different steps of the data engineering lifecycle. This is all to ensure efficient resource utilization and consistency. + +Containers package an application and its dependencies into a portable, isolated unit that runs consistently across environments. Container orchestration automates the deployment, scaling, and management of these containers across a cluster. Together, containers and orchestration form the foundation for running modern data workloads in cloud and hybrid environments. Visit the following resources to learn more: - [@article@What are Containers?](https://cloud.google.com/learn/what-are-containers) - [@article@Containers - The New Stack](https://thenewstack.io/category/containers/) -- [@article@An Introduction to Data Orchestration: Process and Benefits](https://www.datacamp.com/blog/introduction-to-data-orchestration-process-and-benefits) -- [@article@What is Container Orchestration?](https://www.redhat.com/en/topics/containers/what-is-container-orchestration) - [@video@What are Containers?](https://www.youtube.com/playlist?list=PLawsLZMfND4nz-WDBZIj8-nbzGFD4S9oz) - [@video@Why You Need Data Orchestration](https://www.youtube.com/watch?v=ZtlS5-G-gng) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/cosmosdb@goL_GqVVTVxXQMGBw992b.md b/src/data/roadmaps/data-engineer/content/cosmosdb@goL_GqVVTVxXQMGBw992b.md index 7f33acc8d..c4b291d0e 100644 --- a/src/data/roadmaps/data-engineer/content/cosmosdb@goL_GqVVTVxXQMGBw992b.md +++ b/src/data/roadmaps/data-engineer/content/cosmosdb@goL_GqVVTVxXQMGBw992b.md @@ -6,5 +6,4 @@ Visit the following resources to learn more: - [@official@What are Containers?](https://azure.microsoft.com/en-us/products/cosmos-db#FAQ) - [@official@CAzure Cosmos DB - Database for the AI Era](https://learn.microsoft.com/en-us/azure/cosmos-db/introduction) -- [@article@CAzure Cosmos DB: A Global-Scale NoSQL Cloud Database](https://www.datacamp.com/tutorial/azure-cosmos-db) - [@video@What is Azure Cosmos DB?](https://www.youtube.com/watch?v=hBY2YcaIOQM&) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/couchdb@-IesOBWPSIlbgvTjBqHcb.md b/src/data/roadmaps/data-engineer/content/couchdb@-IesOBWPSIlbgvTjBqHcb.md index 638b7d35d..e68ba3d25 100644 --- a/src/data/roadmaps/data-engineer/content/couchdb@-IesOBWPSIlbgvTjBqHcb.md +++ b/src/data/roadmaps/data-engineer/content/couchdb@-IesOBWPSIlbgvTjBqHcb.md @@ -1,9 +1,8 @@ # CouchDB - -Apache CouchDB is an open source NoSQL document database that collects and stores data in JSON-based document formats. Unlike relational databases, CouchDB uses a schema-free data model, which simplifies record management across various computing devices, mobile phones and web browsers. In CouchDB, each document is uniquely named in the database, and CouchDB provides a RESTful HTTP API for reading and updating (add, edit, delete) database documents. Documents are the primary unit of data in CouchDB and consist of any number of fields and attachments. + +Apache CouchDB is an open-source document database that uses JSON for documents and HTTP as its API. It is designed for reliability and offline-first use cases, with a built-in replication protocol that syncs data between devices and servers. CouchDB is used in scenarios where data needs to be available and writable even without a network connection. Visit the following resources to learn more: -- [@official@CouchDB](hhttps://couchdb.apache.org/) - [@official@CouchDB Documentation](https://docs.couchdb.org/en/stable/intro/overview.html) - [@article@What is CouchDB?](https://www.ibm.com/think/topics/couchdb) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/data-analytics@V30v5RLQrWSMBUIsZQG1o.md b/src/data/roadmaps/data-engineer/content/data-analytics@V30v5RLQrWSMBUIsZQG1o.md index 02972ce86..76569169a 100644 --- a/src/data/roadmaps/data-engineer/content/data-analytics@V30v5RLQrWSMBUIsZQG1o.md +++ b/src/data/roadmaps/data-engineer/content/data-analytics@V30v5RLQrWSMBUIsZQG1o.md @@ -1,16 +1,10 @@ # Data Analytics - -Data Analytics involves extracting meaningful insights from raw data to drive decision-making processes. It includes a wide range of techniques and disciplines ranging from the simple data compilation to advanced algorithms and statistical analysis. Data analysts, as ambassadors of this domain, employ these techniques to answer various questions: - -* Descriptive Analytics _(what happened in the past?)_ -* Diagnostic Analytics _(why did it happened in the past?)_ -* Predictive Analytics _(what will happen in the future?)_ -* Prescriptive Analytics _(how can we make it happen?)_ + +Data analytics is the process of examining datasets to draw conclusions and support decision-making. It covers a spectrum from descriptive analytics (what happened) to diagnostic (why it happened), predictive (what might happen), and prescriptive (what to do). Data engineers build the infrastructure that makes analytics possible by ensuring clean, accessible, and timely data. Visit the following resources to learn more: - [@course@Introduction to Data Analytics](https://www.coursera.org/learn/introduction-to-data-analytics) - [@article@The 4 Types of Data Analysis: Ultimate Guide](https://careerfoundry.com/en/blog/data-analytics/different-types-of-data-analysis/) -- [@article@What is Data Analysis? An Expert Guide With Examples](https://www.datacamp.com/blog/what-is-data-analysis-expert-guide) - [@video@Descriptive vs Diagnostic vs Predictive vs Prescriptive Analytics: What's the Difference?](https://www.youtube.com/watch?v=QoEpC7jUb9k) - [@video@Types of Data Analytics](https://www.youtube.com/watch?v=lsZnSgxMwBA) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/data-collection-considerations@wDDWQgMVBYK4WcmHq_d6l.md b/src/data/roadmaps/data-engineer/content/data-collection-considerations@wDDWQgMVBYK4WcmHq_d6l.md index 9acc762cb..da2b92a1e 100644 --- a/src/data/roadmaps/data-engineer/content/data-collection-considerations@wDDWQgMVBYK4WcmHq_d6l.md +++ b/src/data/roadmaps/data-engineer/content/data-collection-considerations@wDDWQgMVBYK4WcmHq_d6l.md @@ -1,12 +1,6 @@ # Data Collection Considerations - -Before designing the technology archecture to collect and store data, you should consider the following factors: - -* **Bounded versus unbounded**. Bounded data has defined start and end points, forming a finite, complete dataset, like the daily sales report. Unbounded data has no predefined limits in time or scope, flowing continuously and potentially indefinitely, such as user interaction events or real-time sensor data. The distinction is critical in data processing, where bounded data is suitable for batch processing, and unbounded data is processed in stream processing or real-time systems. -* **Frequency.** Collection processes can be batch, micro-batch, or real-time, depending on the frequency you need to store the data. -* **Synchronous versus asynchronous.** Synchronous ingestion is a process where the system waits for a response from the data source before proceeding. In contrast, asynchronous ingestion is a process where data is ingested without waiting for a response from the data source. Each approach has its benefits and drawbacks, and the choice depends on the specific requirements of the data ingestion process and the business needs. -* **Throughput and scalability.** As data demands grow, you will need scalable ingestion solutions to keep pace. Scalable data ingestion pipelines ensure that systems can handle increasing data volumes without compromising performance. Without scalable ingestion, data pipelines face challenges like bottlenecks and data loss. Bottlenecks occur when components can't process data fast enough, leading to delays and reduced throughput. Data loss happens when systems are overwhelmed, causing valuable information to be discarded or corrupted. -* **Reliability and durability.** Data reliability in the ingestion phase means ensuring that the acquired data from various sources is accurate, consistent, and trustworthy as it enters the data pipeline. Durability entails making sure that data isn’t lost or corrupted during the data collection process. + +When collecting data, engineers must account for reliability, latency, volume, and schema consistency. Other considerations include data privacy regulations, deduplication, and handling of missing or malformed records. Good collection design reduces problems downstream in the pipeline. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/data-engineering-lifecycle@Ouph2bHeLQsrHl45ar4Cs.md b/src/data/roadmaps/data-engineer/content/data-engineering-lifecycle@Ouph2bHeLQsrHl45ar4Cs.md index 5265218ed..2f1e6a273 100644 --- a/src/data/roadmaps/data-engineer/content/data-engineering-lifecycle@Ouph2bHeLQsrHl45ar4Cs.md +++ b/src/data/roadmaps/data-engineer/content/data-engineering-lifecycle@Ouph2bHeLQsrHl45ar4Cs.md @@ -1,13 +1,6 @@ # Data Engineering Lifecycle - -The data engineering lifecycle encompasses the entire process of transforming raw data into a useful end product. It involves several stages, each with specific roles and responsibilities. This lifecycle ensures that data is handled efficiently and effectively, from its initial generation to its final consumption. - -It involves 4 steps: - -1. Data Generation: Collecting data from various source systems. -2. Data Storage: Safely storing data for future processing and analysis. -3. Data Ingestion: Transforming and bringing data into a centralized system. -4. Data Serving: Providing data to end-users for decision-making and operational purposes. + +The data engineering lifecycle describes the stages data moves through from creation to consumption. These stages typically include generation, ingestion, storage, transformation, and serving. Each stage has its own tools, failure modes, and design considerations. Understanding the full lifecycle helps engineers make better decisions about architecture and tooling. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/data-engineering-lifecycle@w3cfuNC-IdUKA7CEXs0fT.md b/src/data/roadmaps/data-engineer/content/data-engineering-lifecycle@w3cfuNC-IdUKA7CEXs0fT.md index 5265218ed..2f1e6a273 100644 --- a/src/data/roadmaps/data-engineer/content/data-engineering-lifecycle@w3cfuNC-IdUKA7CEXs0fT.md +++ b/src/data/roadmaps/data-engineer/content/data-engineering-lifecycle@w3cfuNC-IdUKA7CEXs0fT.md @@ -1,13 +1,6 @@ # Data Engineering Lifecycle - -The data engineering lifecycle encompasses the entire process of transforming raw data into a useful end product. It involves several stages, each with specific roles and responsibilities. This lifecycle ensures that data is handled efficiently and effectively, from its initial generation to its final consumption. - -It involves 4 steps: - -1. Data Generation: Collecting data from various source systems. -2. Data Storage: Safely storing data for future processing and analysis. -3. Data Ingestion: Transforming and bringing data into a centralized system. -4. Data Serving: Providing data to end-users for decision-making and operational purposes. + +The data engineering lifecycle describes the stages data moves through from creation to consumption. These stages typically include generation, ingestion, storage, transformation, and serving. Each stage has its own tools, failure modes, and design considerations. Understanding the full lifecycle helps engineers make better decisions about architecture and tooling. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/data-engineering-vs-data-science@jJukG4XxfFcID_VlQKqe-.md b/src/data/roadmaps/data-engineer/content/data-engineering-vs-data-science@jJukG4XxfFcID_VlQKqe-.md index f863635bf..d2a23d959 100644 --- a/src/data/roadmaps/data-engineer/content/data-engineering-vs-data-science@jJukG4XxfFcID_VlQKqe-.md +++ b/src/data/roadmaps/data-engineer/content/data-engineering-vs-data-science@jJukG4XxfFcID_VlQKqe-.md @@ -4,5 +4,4 @@ Data engineering and data science are distinct but complementary roles within th Visit the following resources to learn more: -- [@article@Data Scientist vs Data Engineer](https://www.datacamp.com/blog/data-scientist-vs-data-engineer) - [@video@Should You Be a Data Scientist, Analyst or Engineer?](https://www.youtube.com/watch?v=dUnKYhripIE) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/data-fabric@-x3QLMYhC67VJQ6EW6BrJ.md b/src/data/roadmaps/data-engineer/content/data-fabric@-x3QLMYhC67VJQ6EW6BrJ.md index 8b6545e7b..33270d516 100644 --- a/src/data/roadmaps/data-engineer/content/data-fabric@-x3QLMYhC67VJQ6EW6BrJ.md +++ b/src/data/roadmaps/data-engineer/content/data-fabric@-x3QLMYhC67VJQ6EW6BrJ.md @@ -1,6 +1,6 @@ # Data Fabric - -A data fabric is a single environment consisting of a unified architecture with services and technologies running on it that architecture that helps a company manage their data. It enables accessing, ingesting, integrating, and sharing data in a environment where the data can be batched or streamed and be in the cloud or on-prem. The ultimate goal of data fabric is to use all your data to gain better insights into your company and make better business decisions. A data fabric includes building blocks such as data pipeline, data access, data lake, data store, data policy, ingestion framework, and data visualization. These building blocks would be used to build platforms or “products” such as a client data integration platform, data hub, governance framework, and a global semantic layer, giving you centralized governance and standardization + +Data fabric is an architectural concept that aims to provide a unified layer for accessing and managing data across heterogeneous environments, including on-premises and multiple clouds. It uses metadata, automation, and integration patterns to connect disparate data sources. Data fabric focuses on making data discoverable and accessible without requiring it to be moved to a central location. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/data-factory-etl@BNGdJSmrNE90rwPa4JoWj.md b/src/data/roadmaps/data-engineer/content/data-factory-etl@BNGdJSmrNE90rwPa4JoWj.md index a2b3c1197..4adc65409 100644 --- a/src/data/roadmaps/data-engineer/content/data-factory-etl@BNGdJSmrNE90rwPa4JoWj.md +++ b/src/data/roadmaps/data-engineer/content/data-factory-etl@BNGdJSmrNE90rwPa4JoWj.md @@ -6,5 +6,4 @@ Visit the following resources to learn more: - [@course@Microsoft Azure - Data Factory](https://www.coursera.org/learn/microsoft-azure---data-factory) - [@official@What is Azure Data Factory?](https://learn.microsoft.com/en-us/azure/data-factory/introduction) -- [@official@Azure Data Factory Documentation](https://learn.microsoft.com/en-gb/azure/data-factory/) - [@official@Azure Data Factory Documentation](https://learn.microsoft.com/en-gb/azure/data-factory/) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/data-generation@AWf1y87pd1JFW71cZ_iE1.md b/src/data/roadmaps/data-engineer/content/data-generation@AWf1y87pd1JFW71cZ_iE1.md index 62d206795..e48333cf9 100644 --- a/src/data/roadmaps/data-engineer/content/data-generation@AWf1y87pd1JFW71cZ_iE1.md +++ b/src/data/roadmaps/data-engineer/content/data-generation@AWf1y87pd1JFW71cZ_iE1.md @@ -1,10 +1,6 @@ # Data Generation - -Data generation refers to the different ways data is produced and generated. Thanks to progress in computing power and storage, as well as technology breakthrough in sensor technology (for example, IoT devices), the number of these so-called source systems is rapidly growing. Data is created in many ways, both analog and digital. - -**Analog data** refers to continuous, real-world information that is represented by a range of values. It can take on any value within a given range and is often used to describe physical quantities like temperature or sounds. - -By contrast, **digital data** is either created by converting analog data to digital form (eg. images or videos) or is the native product of a digital system, such as logs from a mobile app or syntetic data. + +Data generation refers to how raw data is produced and originates in a system. Data can come from user interactions, application logs, IoT sensors, databases, APIs, and many other sources. Understanding where data comes from and how it is structured at the source is the starting point for any data pipeline design. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/data-hub@OiWleAdMbPtisrJpk2eSJ.md b/src/data/roadmaps/data-engineer/content/data-hub@OiWleAdMbPtisrJpk2eSJ.md index 511039bc5..d1bf68893 100644 --- a/src/data/roadmaps/data-engineer/content/data-hub@OiWleAdMbPtisrJpk2eSJ.md +++ b/src/data/roadmaps/data-engineer/content/data-hub@OiWleAdMbPtisrJpk2eSJ.md @@ -1,8 +1,6 @@ # Data Hub - -A **data hub** is an architecture that provides a central point for the flow of data between multiple sources and applications, enabling organizations to collect, integrate, and manage data efficiently. Unlike traditional data storage solutions, a data hub’s purpose focuses on data integration and accessibility. The design supports real-time data exchange, which makes accessing, analyzing, and acting on the data faster and easier. - -A data hub differs from a data warehouse in that it is generally unintegrated and often at different grains. It differs from an operational data store because a data hub does not need to be limited to operational data. A data hub differs from a data lake by homogenizing data and possibly serving data in multiple desired formats, rather than simply storing it in one place, and by adding other value to the data such as de-duplication, quality, security, and a standardized set of query services. + +A data hub is a centralized platform that acts as an integration point for data flowing between multiple systems. Unlike a data warehouse, a data hub focuses on data movement and integration rather than storage for analytics. It often combines features of a message broker, metadata catalog, and integration layer. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/data-ingestion@CvCOkyWcgzaUJec_v5F4L.md b/src/data/roadmaps/data-engineer/content/data-ingestion@CvCOkyWcgzaUJec_v5F4L.md index e451d2d61..1833ea345 100644 --- a/src/data/roadmaps/data-engineer/content/data-ingestion@CvCOkyWcgzaUJec_v5F4L.md +++ b/src/data/roadmaps/data-engineer/content/data-ingestion@CvCOkyWcgzaUJec_v5F4L.md @@ -5,4 +5,4 @@ Data ingestion is the third step in the data engineering lifecycle. It entails t Visit the following resources to learn more: - [@article@What is Data Ingestion?](https://www.ibm.com/think/topics/data-ingestion) -- [@article@WData Ingestion](https://www.qlik.com/us/data-ingestion) \ No newline at end of file +- [@article@Data Ingestion](https://www.qlik.com/us/data-ingestion) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/data-lake@y0Lxz_wVyQ6lr1hvCsufa.md b/src/data/roadmaps/data-engineer/content/data-lake@y0Lxz_wVyQ6lr1hvCsufa.md index bec07f16b..9d573d733 100644 --- a/src/data/roadmaps/data-engineer/content/data-lake@y0Lxz_wVyQ6lr1hvCsufa.md +++ b/src/data/roadmaps/data-engineer/content/data-lake@y0Lxz_wVyQ6lr1hvCsufa.md @@ -1,6 +1,6 @@ -# Data lakes - -**Data Lakes** are large-scale data repository systems that store raw, untransformed data, in various formats, from multiple sources. They're often used for big data and real-time analytics requirements. Data lakes preserve the original data format and schema which can be modified as necessary. +# Data Lake + +A data lake is a centralized storage repository that holds large amounts of raw data in its native format, including structured, semi-structured, and unstructured data. Unlike a data warehouse, a data lake does not enforce a schema on ingestion. Data is stored cheaply at scale and processed when needed, which enables flexibility for future analysis. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/data-lineage@pKewO7Ef3GBXL4MDK62QG.md b/src/data/roadmaps/data-engineer/content/data-lineage@pKewO7Ef3GBXL4MDK62QG.md index 66905f59e..c0fca2209 100644 --- a/src/data/roadmaps/data-engineer/content/data-lineage@pKewO7Ef3GBXL4MDK62QG.md +++ b/src/data/roadmaps/data-engineer/content/data-lineage@pKewO7Ef3GBXL4MDK62QG.md @@ -4,5 +4,4 @@ Visit the following resources to learn more: -- [@article@What is Data Lineage? - IBM](https://www.ibm.com/topics/data-lineage) -- [@article@What is Data Lineage? - Datacamp](https://www.datacamp.com/blog/data-lineage) \ No newline at end of file +- [@article@What is Data Lineage? - IBM](https://www.ibm.com/topics/data-lineage) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/data-mesh@D7qtosIbsQuIY3OWl_Hwc.md b/src/data/roadmaps/data-engineer/content/data-mesh@D7qtosIbsQuIY3OWl_Hwc.md index 99ccce9ad..4bbc0cad4 100644 --- a/src/data/roadmaps/data-engineer/content/data-mesh@D7qtosIbsQuIY3OWl_Hwc.md +++ b/src/data/roadmaps/data-engineer/content/data-mesh@D7qtosIbsQuIY3OWl_Hwc.md @@ -5,5 +5,4 @@ A data mesh is a modern approach to data architecture that shifts data managemen Visit the following resources to learn more: - [@article@What Is a Data Mesh? - AWS](https://aws.amazon.com/what-is/data-mesh) -- [@article@What Is a Data Mesh? - Datacamp](https://www.datacamp.com/blog/data-mesh) - [@video@Data Mesh Architecture](https://www.datamesh-architecture.com/) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/data-modelling-techniques@SlQHO8n97F7-_fc6EUXlj.md b/src/data/roadmaps/data-engineer/content/data-modelling-techniques@SlQHO8n97F7-_fc6EUXlj.md index e0771ba52..834cbef18 100644 --- a/src/data/roadmaps/data-engineer/content/data-modelling-techniques@SlQHO8n97F7-_fc6EUXlj.md +++ b/src/data/roadmaps/data-engineer/content/data-modelling-techniques@SlQHO8n97F7-_fc6EUXlj.md @@ -1,13 +1,7 @@ # Data Modelling Techniques - -A data model is a specification of data structures and business rules. It creates a visual representation of data and illustrates how different data elements are related to each other. Different techniques are employed depending on the complexity of the data and the goals. Below you can find a list with the most common data modelling techniques: - -* **Entity-relationship modeling.** It's one of the most common techniques used to represent data. It's based on three elements: Entities (objects or things within the system), relationships (how these entities interact with each other), and attributes (properties of the entities). -* **Dimensional modeling.** Dimensional modeling is widely used in data warehousing and analytics, where data is often represented in terms of facts and dimensions. This technique simplifies complex data by organizing it into a star or snowflake schema. -* **Object-oriented modeling.** Object-oriented modeling is used to represent complex systems, where data and the functions that operate on it are encapsulated as objects. This technique is preferred for modeling applications with complex, interrelated data and behaviors -* **NoSQL modeling.** NoSQL modeling techniques are designed for flexible, schema-less databases. These approaches are often used when data structures are less rigid or evolve over time + +Data modelling is the process of defining how data is structured and related within a storage system. Common techniques include entity-relationship (ER) modelling for transactional databases and dimensional modelling (star and snowflake schemas) for analytics. The choice of model affects query performance, flexibility, and how easy it is to evolve the schema over time. Visit the following resources to learn more: -- [@article@7 data modeling techniques and concepts for business](https://www.techtarget.com/searchdatamanagement/tip/7-data-modeling-techniques-and-concepts-for-business) -- [@article@@articleData Modeling Explained: Techniques, Examples, and Best Practices](https://www.datacamp.com/blog/data-modeling) \ No newline at end of file +- [@article@7 data modeling techniques and concepts for business](https://www.techtarget.com/searchdatamanagement/tip/7-data-modeling-techniques-and-concepts-for-business) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/data-normalization@kVPEoUX-ZAGwstieD20Qa.md b/src/data/roadmaps/data-engineer/content/data-normalization@kVPEoUX-ZAGwstieD20Qa.md index 5ac51d396..8db14169b 100644 --- a/src/data/roadmaps/data-engineer/content/data-normalization@kVPEoUX-ZAGwstieD20Qa.md +++ b/src/data/roadmaps/data-engineer/content/data-normalization@kVPEoUX-ZAGwstieD20Qa.md @@ -1,9 +1,8 @@ -# Database Normalization - -Database normalization is the process of structuring a relational database in accordance with a series of so-called normal forms in order to reduce data redundancy and improve data integrity. It was first proposed by Edgar F. Codd as part of his relational model. Normalization entails organizing the columns (attributes) and tables (relations) of a database to ensure that their dependencies are properly enforced by database integrity constraints. It is accomplished by applying some formal rules either by a process of synthesis (creating a new database design) or decomposition (improving an existing database design). +# Data Normalization + +Data normalization is the process of organizing a relational database to reduce redundancy and improve data integrity. It involves decomposing tables into smaller, related ones according to normal forms (1NF, 2NF, 3NF, etc.). Normalized schemas are easier to maintain but may require more joins when querying. Visit the following resources to learn more: - [@article@What is Normalization in DBMS (SQL)? 1NF, 2NF, 3NF, BCNF Database with Example](https://www.guru99.com/database-normalization.html) -- [@video@Complete guide to Database Normalization in SQL](https://www.youtube.com/watch?v=rBPQ5fg_kiY) -- [@feed@Explore top posts about Database](https://app.daily.dev/tags/database?ref=roadmapsh) \ No newline at end of file +- [@video@Complete guide to Database Normalization in SQL](https://www.youtube.com/watch?v=rBPQ5fg_kiY) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/data-quality@cStrYgFZA2NuYq8TdWWP_.md b/src/data/roadmaps/data-engineer/content/data-quality@cStrYgFZA2NuYq8TdWWP_.md index 7efcc6038..80177c6d6 100644 --- a/src/data/roadmaps/data-engineer/content/data-quality@cStrYgFZA2NuYq8TdWWP_.md +++ b/src/data/roadmaps/data-engineer/content/data-quality@cStrYgFZA2NuYq8TdWWP_.md @@ -1,5 +1,3 @@ # Data Quality - -Ensuring quality involves validating the accuracy, completeness, consistency, and reliability of the data collected from each source. The fact that you do it from one source or multiple is almost irrelevant since the only extra task would be to homogenize the final schema of the data, ensuring deduplication and normalization. - -This last part typically includes verifying the credibility of each data source, standardizing formats (like date/time or currency), performing schema alignment, and running profiling to detect anomalies, duplicates, or mismatches before integrating the data for analysis. \ No newline at end of file + +Data quality refers to how well data meets the requirements for its intended use in terms of accuracy, completeness, consistency, timeliness, and validity. Poor data quality leads to incorrect analysis and poor decisions. Data engineers implement quality checks at ingestion and transformation stages to catch and prevent data issues. \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/data-storage@wydtifF3ZhMWCbVt8Hd2t.md b/src/data/roadmaps/data-engineer/content/data-storage@wydtifF3ZhMWCbVt8Hd2t.md index d2c79875d..1c6a6f2ae 100644 --- a/src/data/roadmaps/data-engineer/content/data-storage@wydtifF3ZhMWCbVt8Hd2t.md +++ b/src/data/roadmaps/data-engineer/content/data-storage@wydtifF3ZhMWCbVt8Hd2t.md @@ -1,6 +1,6 @@ # Data Storage - -Data storage is the process of saving and preserving digital information on various physical or cloud-based media for future retrieval and use. It encompasses the use of technologies and devices like hard drives and cloud platforms to store data. + +Data storage in the engineering lifecycle refers to where and how data is persisted after it is generated or ingested. The choice of storage system depends on access patterns, data volume, latency requirements, and cost. Options range from relational databases to object storage, data lakes, and columnar warehouses. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/data-structures-and-algorithms@fqmn6DPOA5MH7UWYv6ayn.md b/src/data/roadmaps/data-engineer/content/data-structures-and-algorithms@fqmn6DPOA5MH7UWYv6ayn.md index 0f0cd602c..8cfa8651b 100644 --- a/src/data/roadmaps/data-engineer/content/data-structures-and-algorithms@fqmn6DPOA5MH7UWYv6ayn.md +++ b/src/data/roadmaps/data-engineer/content/data-structures-and-algorithms@fqmn6DPOA5MH7UWYv6ayn.md @@ -1,12 +1,10 @@ -# DataStructures and Algorithms - -**Data Structures** are primarily used to collect, organize and perform operations on the stored data more effectively. They are essential for designing advanced-level Android applications. Examples include Array, Linked List, Stack, Queue, Hash Map, and Tree. - -**Algorithms** are a sequence of instructions or rules for performing a particular task. Algorithms can be used for data searching, sorting, or performing complex business logic. Some commonly used algorithms are Binary Search, Bubble Sort, Selection Sort, etc. A deep understanding of data structures and algorithms is crucial in optimizing the performance and the memory consumption of data pipelines +# Data Structures and Algorithms + +Data structures and algorithms form the foundation for writing efficient code. This knowledge is relevant for data engineers when optimizing queries, designing storage schemas, and building processing logic that scales. Common topics include arrays, hash maps, trees, sorting, and complexity analysis. Visit the following resources to learn more: +- [@roadmap@Visit the Dedicated DSA Roadmap](https://roadmap.sh/datastructures-and-algorithms) - [@article@Interview Questions about Data Structures](https://www.csharpstar.com/csharp-algorithms/) - [@video@Data Structures Illustrated](https://www.youtube.com/watch?v=9rhT3P1MDHk&list=PLkZYeFmDuaN2-KUIv-mvbjfKszIGJ4FaY) -- [@video@Intro to Algorithms](https://www.youtube.com/watch?v=rL8X2mlNHPM) -- [@feed@Explore top posts about Algorithms](https://app.daily.dev/tags/algorithms?ref=roadmapsh) \ No newline at end of file +- [@video@Intro to Algorithms](https://www.youtube.com/watch?v=rL8X2mlNHPM) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/data-warehouse@ArOoKuf9scAURs8NRjAru.md b/src/data/roadmaps/data-engineer/content/data-warehouse@ArOoKuf9scAURs8NRjAru.md index 0bee40913..012821a89 100644 --- a/src/data/roadmaps/data-engineer/content/data-warehouse@ArOoKuf9scAURs8NRjAru.md +++ b/src/data/roadmaps/data-engineer/content/data-warehouse@ArOoKuf9scAURs8NRjAru.md @@ -1,8 +1,8 @@ # Data Warehouse - -**Data Warehouses** are data storage systems which are designed for analyzing, reporting and integrating with transactional systems. The data in a warehouse is clean, consistent, and often transformed to meet wide-range of business requirements. Hence, data warehouses provide structured data but require more processing and management compared to data lakes. + +A data warehouse stores structured, processed data from operational systems, optimized for analytical queries. It typically uses columnar storage and is populated through ETL or ELT processes. Common cloud data warehouses include Google BigQuery, Snowflake, and Amazon Redshift. Visit the following resources to learn more: - [@article@What Is a Data Warehouse?](https://www.oracle.com/database/what-is-a-data-warehouse/) -- [@video@@hat is a Data Warehouse?](https://www.youtube.com/watch?v=k4tK2ttdSDg) \ No newline at end of file +- [@video@What is a Data Warehouse?](https://www.youtube.com/watch?v=k4tK2ttdSDg) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/data-warehousing-architectures@J854xPM1X0BWlhtJw7Hs_.md b/src/data/roadmaps/data-engineer/content/data-warehousing-architectures@J854xPM1X0BWlhtJw7Hs_.md index d3cbd7d6e..e076ff87c 100644 --- a/src/data/roadmaps/data-engineer/content/data-warehousing-architectures@J854xPM1X0BWlhtJw7Hs_.md +++ b/src/data/roadmaps/data-engineer/content/data-warehousing-architectures@J854xPM1X0BWlhtJw7Hs_.md @@ -1,3 +1,3 @@ # Data Warehousing Architectures - -Data Warehousing Architectures refers to the different systems and solutions for storing data. Options include traditional data warehouse, data marts, data lakes and data mesh architectures. \ No newline at end of file + +Data warehousing architectures describe how data is organized, stored, and accessed across a warehouse system. Common patterns include traditional ETL-based warehouses, cloud-native warehouses, data lakehouse architectures, and federated query systems. The choice of architecture affects cost, query performance, scalability, and how fresh the data available for analysis is. \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/database-fundamentals@g4UC0go7OPCJYJlac9w-i.md b/src/data/roadmaps/data-engineer/content/database-fundamentals@g4UC0go7OPCJYJlac9w-i.md index 3d8206a4c..b939c5558 100644 --- a/src/data/roadmaps/data-engineer/content/database-fundamentals@g4UC0go7OPCJYJlac9w-i.md +++ b/src/data/roadmaps/data-engineer/content/database-fundamentals@g4UC0go7OPCJYJlac9w-i.md @@ -1,17 +1,10 @@ -# Database fundamentals - -A database is a collection of useful data of one or more related organizations structured in a way to make data an asset to the organization. A database management system is a software designed to assist in maintaining and extracting large collections of data in a timely fashion. - -A **Relational database** is a type of database that stores and provides access to data points that are related to one another. Relational databases store data in a series of tables. - -**NoSQL databases** offer data storage and retrieval that is modelled differently to "traditional" relational databases. NoSQL databases typically focus more on horizontal scaling, eventual consistency, speed and flexibility and is used commonly for big data and real-time streaming applications. +# Database Fundamentals + +Database fundamentals cover the core concepts that apply across most relational database systems: how data is organized into tables, how queries are executed, and how the database ensures consistency and durability. Topics include normalization, indexing, transactions, and query optimization. These concepts apply whether using PostgreSQL, MySQL, or any other relational system. Visit the following resources to learn more: - [@article@Oracle: What is a Database?](https://www.oracle.com/database/what-is-database/) -- [@article@Prisma.io: What are Databases?](https://www.prisma.io/dataguide/intro/what-are-databases) -- [@article@Intro To Relational Databases](https://www.udacity.com/course/intro-to-relational-databases--ud197) - [@article@NoSQL Explained](https://www.mongodb.com/nosql-explained) - [@video@What is Relational Database](https://youtu.be/OqjJjpjDRLc) -- [@video@How do NoSQL Databases work](https://www.youtube.com/watch?v=0buKQHokLK8) -- [@feed@Explore top posts about Database](https://app.daily.dev/tags/database?ref=roadmapsh) \ No newline at end of file +- [@video@How do NoSQL Databases work](https://www.youtube.com/watch?v=0buKQHokLK8) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/databricks-delta-lake@fhfyoWekmYvEs-jdP2mJo.md b/src/data/roadmaps/data-engineer/content/databricks-delta-lake@fhfyoWekmYvEs-jdP2mJo.md index c2f030d70..d72d7e7c0 100644 --- a/src/data/roadmaps/data-engineer/content/databricks-delta-lake@fhfyoWekmYvEs-jdP2mJo.md +++ b/src/data/roadmaps/data-engineer/content/databricks-delta-lake@fhfyoWekmYvEs-jdP2mJo.md @@ -6,5 +6,4 @@ Visit the following resources to learn more: - [@book@The Delta Lake Series — Fundamentals and Performance](https://www.databricks.com/resources/ebook/the-delta-lake-series-fundamentals-performance) - [@official@What is Delta Lake in Databricks?](https://docs.databricks.com/aws/en/delta) -- [@article@Delta Table in Databricks: A Complete Guide](https://www.datacamp.com/tutorial/delta-table-in-databricks) - [@video@Delta Lake](https://www.databricks.com/resources/demos/videos/lakehouse-platform/delta-lake) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/datadog@Zoa4JEGrSKjVwUNer4Go1.md b/src/data/roadmaps/data-engineer/content/datadog@Zoa4JEGrSKjVwUNer4Go1.md index 7ecc60051..c293f577a 100644 --- a/src/data/roadmaps/data-engineer/content/datadog@Zoa4JEGrSKjVwUNer4Go1.md +++ b/src/data/roadmaps/data-engineer/content/datadog@Zoa4JEGrSKjVwUNer4Go1.md @@ -4,5 +4,4 @@ Datadog is a monitoring and analytics platform for large-scale applications. It Visit the following resources to learn more: -- [@official@Datadog](https://www.datadoghq.com/) - [@official@Datadog Documentation](https://docs.datadoghq.com/) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/dbt@SgYLIkMtLVPlw8Qo5j0Fb.md b/src/data/roadmaps/data-engineer/content/dbt@SgYLIkMtLVPlw8Qo5j0Fb.md index b20ac3266..c1f0cc8d6 100644 --- a/src/data/roadmaps/data-engineer/content/dbt@SgYLIkMtLVPlw8Qo5j0Fb.md +++ b/src/data/roadmaps/data-engineer/content/dbt@SgYLIkMtLVPlw8Qo5j0Fb.md @@ -5,5 +5,4 @@ dbt, also known as the data build tool, is designed to simplify the management o Visit the following resources to learn more: - [@course@dbt Official Courses](https://learn.getdbt.com/catalog) -- [@official@dbt](https://www.getdbt.com/product/what-is-dbt) - [@official@dbt Documentation](https://docs.getdbt.com/docs/build/documentation) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/distributed-systems-basics@c1dadtQgbqXwcsQhI6de0.md b/src/data/roadmaps/data-engineer/content/distributed-systems-basics@c1dadtQgbqXwcsQhI6de0.md index 4ca844605..09d2083b7 100644 --- a/src/data/roadmaps/data-engineer/content/distributed-systems-basics@c1dadtQgbqXwcsQhI6de0.md +++ b/src/data/roadmaps/data-engineer/content/distributed-systems-basics@c1dadtQgbqXwcsQhI6de0.md @@ -6,4 +6,4 @@ Visit the following resources to learn more: - [@article@Introduction to Distributed Systems](https://www.freecodecamp.org/news/a-thorough-introduction-to-distributed-systems-3b91562c9b3c/) - [@article@Distributed Systems Guide](https://www.baeldung.com/cs/distributed-systems-guide) -- [@video@Quick overview](https://www.youtube.com/watch?v=IJWwfMyPu1c) \ No newline at end of file +- [@video@Distributed Systems Explained](https://www.youtube.com/watch?v=IJWwfMyPu1c) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/docker@OQ3RqVgWEMxpAtrrjOG5U.md b/src/data/roadmaps/data-engineer/content/docker@OQ3RqVgWEMxpAtrrjOG5U.md index 4c04ae9ae..ad1683d96 100644 --- a/src/data/roadmaps/data-engineer/content/docker@OQ3RqVgWEMxpAtrrjOG5U.md +++ b/src/data/roadmaps/data-engineer/content/docker@OQ3RqVgWEMxpAtrrjOG5U.md @@ -1,11 +1,10 @@ # Docker - -Docker is an open-source platform that automates the deployment, scaling, and management of applications using containerization technology. It enables developers to package applications with all their dependencies into standardized units called containers, ensuring consistent behavior across different environments. Docker provides a lightweight alternative to full machine virtualization, using OS-level virtualization to run multiple isolated systems on a single host. Its ecosystem includes tools for building, sharing, and running containers, such as Docker Engine, Docker Hub, and Docker Compose. Docker has become integral to modern DevOps practices, facilitating microservices architectures, continuous integration/deployment pipelines, and efficient resource utilization in both development and production environments. + +Docker is the most widely used platform for building, shipping, and running containers. It packages code and its dependencies into a lightweight, portable image that runs the same in any environment. Data engineers use Docker to containerize pipeline code, ensure reproducible environments, and simplify deployment. Visit the following resources to learn more: - [@roadmap@Visit Dedicated Docker Roadmap](https://roadmap.sh/docker) - [@official@Docker Documentation](https://docs.docker.com/) - [@video@Docker Tutorial](https://www.youtube.com/watch?v=RqTEHSBrYFw) -- [@video@Docker simplified in 55 seconds](https://youtu.be/vP_4DlOH1G4) -- [@feed@Explore top posts about Docker](https://app.daily.dev/tags/docker?ref=roadmapsh) \ No newline at end of file +- [@video@Docker simplified in 55 seconds](https://youtu.be/vP_4DlOH1G4) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/document@sGkAOVl3C-xIIAdtDH9jq.md b/src/data/roadmaps/data-engineer/content/document@sGkAOVl3C-xIIAdtDH9jq.md index 72a817649..c590b08fc 100644 --- a/src/data/roadmaps/data-engineer/content/document@sGkAOVl3C-xIIAdtDH9jq.md +++ b/src/data/roadmaps/data-engineer/content/document@sGkAOVl3C-xIIAdtDH9jq.md @@ -5,4 +5,4 @@ Visit the following resources to learn more: - [@article@What is a Document Database?](https://www.mongodb.com/resources/basics/databases/document-databases) -- [@article@HDocument-oriented database](https://en.wikipedia.org/wiki/Document-oriented_database) \ No newline at end of file +- [@article@Document-oriented database](https://en.wikipedia.org/wiki/Document-oriented_database) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/dynamodb@BDfpCDOxXZ-Tp0Abj_CVW.md b/src/data/roadmaps/data-engineer/content/dynamodb@BDfpCDOxXZ-Tp0Abj_CVW.md index 0b048ba8c..bdccba255 100644 --- a/src/data/roadmaps/data-engineer/content/dynamodb@BDfpCDOxXZ-Tp0Abj_CVW.md +++ b/src/data/roadmaps/data-engineer/content/dynamodb@BDfpCDOxXZ-Tp0Abj_CVW.md @@ -1,6 +1,6 @@ # DynamoDB - -Amazon DynamoDB is a fully managed NoSQL database solution that provides fast and predictable performance with seamless scalability. It is a key-value and document database that delivers single-digit millisecond performance at any scale. DynamoDB can handle more than 10 trillion requests per day and support peaks of more than 20 million requests per second. It maintains high durability of data via automatic replication across three different zones in an Amazon defined region. + +Amazon DynamoDB is a fully managed key-value and document database service on AWS. It provides single-digit millisecond performance at any scale and handles replication and scaling automatically. DynamoDB is commonly used for applications that require predictable performance and high availability without database administration. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/ecpa@g1VwuSupohuDAT2O4hTXx.md b/src/data/roadmaps/data-engineer/content/ecpa@g1VwuSupohuDAT2O4hTXx.md index 6b052f6ae..e80da8643 100644 --- a/src/data/roadmaps/data-engineer/content/ecpa@g1VwuSupohuDAT2O4hTXx.md +++ b/src/data/roadmaps/data-engineer/content/ecpa@g1VwuSupohuDAT2O4hTXx.md @@ -1,6 +1,6 @@ # ECPA - -The California Consumer Privacy Act (CCPA) is a California state law enacted in 2020 that protects and enforces the rights of Californians regarding the privacy of consumers’ personal information (PI). + +The Electronic Communications Privacy Act (ECPA) is a US federal law that regulates government access to electronic communications and stored data. It sets rules for when law enforcement can intercept communications or compel disclosure of stored data from service providers. Data engineers working with communication data must be aware of ECPA requirements when designing storage and access controls. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/elasticsearch@_F53cV3ln2yu0ics5BFfx.md b/src/data/roadmaps/data-engineer/content/elasticsearch@_F53cV3ln2yu0ics5BFfx.md index a49237c83..c07f69e44 100644 --- a/src/data/roadmaps/data-engineer/content/elasticsearch@_F53cV3ln2yu0ics5BFfx.md +++ b/src/data/roadmaps/data-engineer/content/elasticsearch@_F53cV3ln2yu0ics5BFfx.md @@ -1,10 +1,10 @@ # Elasticsearch - -Elastic search at its core is a document-oriented search engine. It is a document based database that lets you INSERT, DELETE , RETRIEVE and even perform analytics on the saved records. But, Elastic Search is unlike any other general purpose database you have worked with, in the past. It's essentially a search engine and offers an arsenal of features you can use to retrieve the data stored in it, as per your search criteria. And that too, at lightning speeds. + +Elasticsearch is a distributed search and analytics engine built on Apache Lucene. It is designed for full-text search, log analysis, and real-time data exploration. Elasticsearch is commonly used as the backend for search features in applications and as a centralized store for log and event data, often alongside Kibana. Visit the following resources to learn more: +- [@roadmap@Visit the Dedicated Elasticsearch Roadmap](https://roadmap.sh/elasticsearch) - [@official@Elasticsearch Website](https://www.elastic.co/elasticsearch/) - [@official@Elasticsearch Documentation](https://www.elastic.co/guide/index.html) -- [@video@What is Elasticsearch](https://www.youtube.com/watch?v=ZP0NmfyfsoM) -- [@feed@Explore top posts about ELK](https://app.daily.dev/tags/elk?ref=roadmapsh) \ No newline at end of file +- [@video@What is Elasticsearch](https://www.youtube.com/watch?v=ZP0NmfyfsoM) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/etl-vs-reverse-etl@LMFREK9dH_7qzx_s2xCjI.md b/src/data/roadmaps/data-engineer/content/etl-vs-reverse-etl@LMFREK9dH_7qzx_s2xCjI.md index 5a22ec40f..2a0ae2f62 100644 --- a/src/data/roadmaps/data-engineer/content/etl-vs-reverse-etl@LMFREK9dH_7qzx_s2xCjI.md +++ b/src/data/roadmaps/data-engineer/content/etl-vs-reverse-etl@LMFREK9dH_7qzx_s2xCjI.md @@ -1,8 +1,6 @@ # ETL vs Reverse ETL - -ETL (Extract, Transform, Load) is a key process in data warehousing, enabling the integration of data from multiple sources into a centralized database. - -Reverse ETL emerged as organizations recognized that their carefully curated data warehouses, while excellent for analysis, created a new form of data silo that prevented operational teams from accessing valuable insights. This methodology addresses the critical gap between analytical insights and operational execution by systematically moving processed data from centralized repositories back to the operational systems where business teams interact with customers and manage daily operations. + +ETL (Extract, Transform, Load) moves data from operational systems into a data warehouse for analysis. Reverse ETL goes in the opposite direction, syncing processed data from the warehouse back into operational tools like Salesforce, HubSpot, or Intercom. The two patterns are complementary and together form a complete data activation workflow. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/functional-testing@E4ND5XaMDGDLtlV7wTzi6.md b/src/data/roadmaps/data-engineer/content/functional-testing@E4ND5XaMDGDLtlV7wTzi6.md index 30c93fe43..e1c2740de 100644 --- a/src/data/roadmaps/data-engineer/content/functional-testing@E4ND5XaMDGDLtlV7wTzi6.md +++ b/src/data/roadmaps/data-engineer/content/functional-testing@E4ND5XaMDGDLtlV7wTzi6.md @@ -5,5 +5,4 @@ Functional testing is a type of software testing that validates the software sys Visit the following resources to learn more: - [@article@What is Functional Testing? Types & Examples](https://www.guru99.com/functional-testing.html) -- [@article@Functional Testing : A Detailed Guide](https://www.browserstack.com/guide/functional-testing) -- [@feed@Explore top posts about Testing](https://app.daily.dev/tags/testing?ref=roadmapsh) \ No newline at end of file +- [@article@Functional Testing : A Detailed Guide](https://www.browserstack.com/guide/functional-testing) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/gdpr@MuPHohc7mJzcH5QdJ-K46.md b/src/data/roadmaps/data-engineer/content/gdpr@MuPHohc7mJzcH5QdJ-K46.md index a4241a0de..2f85c6286 100644 --- a/src/data/roadmaps/data-engineer/content/gdpr@MuPHohc7mJzcH5QdJ-K46.md +++ b/src/data/roadmaps/data-engineer/content/gdpr@MuPHohc7mJzcH5QdJ-K46.md @@ -1,6 +1,6 @@ -# GDPR in API Design - -The General Data Protection Regulation (GDPR) is an essential standard in API Design that addresses the storage, transfer, and processing of personal data of individuals within the European Union. With regards to API Design, considerations must be given on how APIs handle, process, and secure the data to conform with GDPR's demands on data privacy and security. This includes requirements for explicit consent, right to erasure, data portability, and privacy by design. Non-compliance with these standards not only leads to hefty fines but may also erode trust from users and clients. As such, understanding the impact and integration of GDPR within API design is pivotal for organizations handling EU residents' data. +# GDPR + +GDPR (General Data Protection Regulation) is a data privacy law enacted by the European Union that governs how personal data of EU residents is collected, stored, processed, and shared. It grants individuals rights over their data, including the right to access, correct, and delete it. Data engineers must design systems that support these rights and comply with GDPR requirements such as data minimization and purpose limitation. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/git-and-github@02TADW_PPVtTU_rWV3jf1.md b/src/data/roadmaps/data-engineer/content/git-and-github@02TADW_PPVtTU_rWV3jf1.md index c0f1e2b07..f6c17aee8 100644 --- a/src/data/roadmaps/data-engineer/content/git-and-github@02TADW_PPVtTU_rWV3jf1.md +++ b/src/data/roadmaps/data-engineer/content/git-and-github@02TADW_PPVtTU_rWV3jf1.md @@ -1,15 +1,10 @@ # Git and GitHub - -**Git** is a free and open source distributed version control system designed to handle everything from small to very large projects with speed and efficiency. - -**GitHub** is a web-based platform that provides hosting for software development and version control using Git. It is widely used by developers and organizations around the world to manage and collaborate on software projects. + +Git is a distributed version control system that tracks changes to code over time. GitHub is a platform built on top of Git that adds collaboration features like pull requests, code review, and CI/CD integrations. Data engineers use Git to manage pipeline code, infrastructure configurations, and shared scripts across teams. Visit the following resources to learn more: - [@roadmap@Visit Dedicated Git & GitHub Roadmap](https://roadmap.sh/git-github) - [@official@Git Documentation](https://git-scm.com/) - [@official@GitHub Documentation](https://docs.github.com/en/get-started/quickstart) -- [@article@Learn Git with Tutorials, News and Tips - Atlassian](https://www.atlassian.com/git) -- [@article@Git Cheat Sheet](https://cs.fyi/guide/git-cheatsheet) -- [@video@What is GitHub?](https://www.youtube.com/watch?v=w3jLJU7DT5E) - [@video@Git & GitHub Crash Course For Beginners](https://www.youtube.com/watch?v=SWYqp7iY_Tc) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/github-actions@N8vpCfSdZCADwO_qceWBK.md b/src/data/roadmaps/data-engineer/content/github-actions@N8vpCfSdZCADwO_qceWBK.md index aea8380be..d463b162b 100644 --- a/src/data/roadmaps/data-engineer/content/github-actions@N8vpCfSdZCADwO_qceWBK.md +++ b/src/data/roadmaps/data-engineer/content/github-actions@N8vpCfSdZCADwO_qceWBK.md @@ -1,6 +1,6 @@ # GitHub Actions - -GitHub Actions is a CI/CD tool integrated directly into GitHub, allowing developers to automate workflows, such as building, testing, and deploying code directly from their repositories. It uses YAML files to define workflows, which can be triggered by various events like pushes, pull requests, or on a schedule. GitHub Actions supports a wide range of actions and integrations, making it highly customizable for different project needs. It provides a marketplace with reusable workflows and actions contributed by the community. With its seamless integration with GitHub, developers can take advantage of features like matrix builds, secrets management, and environment-specific configurations to streamline and enhance their development and deployment processes. + +GitHub Actions is a CI/CD platform built into GitHub. It allows developers to define automated workflows as YAML files that trigger on code events like pushes and pull requests. GitHub Actions is widely used to run tests, lint code, build Docker images, and deploy data pipelines. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/gitlab-ci@IYIO4S3DO5xkLD__XT5Dp.md b/src/data/roadmaps/data-engineer/content/gitlab-ci@IYIO4S3DO5xkLD__XT5Dp.md index 562dedaf9..d2e873fc1 100644 --- a/src/data/roadmaps/data-engineer/content/gitlab-ci@IYIO4S3DO5xkLD__XT5Dp.md +++ b/src/data/roadmaps/data-engineer/content/gitlab-ci@IYIO4S3DO5xkLD__XT5Dp.md @@ -4,9 +4,7 @@ GitLab offers a CI/CD service that can be used as a SaaS offering or self-manage Visit the following resources to learn more: -- [@official@GitLab](https://gitlab.com/) - [@official@GitLab Documentation](https://docs.gitlab.com/) - [@official@Get Started with GitLab CI](https://docs.gitlab.com/ee/ci/quick_start/) - [@official@Learn GitLab Tutorials](https://docs.gitlab.com/ee/tutorials/) -- [@official@GitLab CI/CD Examples](https://docs.gitlab.com/ee/ci/examples/) -- [@feed@Explore top posts about GitLab](https://app.daily.dev/tags/gitlab?ref=roadmapsh) \ No newline at end of file +- [@official@GitLab CI/CD Examples](https://docs.gitlab.com/ee/ci/examples/) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/glue-etl@nD36-PXHzOXePM7j9u_O_.md b/src/data/roadmaps/data-engineer/content/glue-etl@nD36-PXHzOXePM7j9u_O_.md index 1f5789b53..481d10580 100644 --- a/src/data/roadmaps/data-engineer/content/glue-etl@nD36-PXHzOXePM7j9u_O_.md +++ b/src/data/roadmaps/data-engineer/content/glue-etl@nD36-PXHzOXePM7j9u_O_.md @@ -1,6 +1,6 @@ -# Amazon RDS (Database) - -Amazon RDS (Relational Database Service) is a web service from Amazon Web Services. It's designed to simplify the setup, operation, and scaling of relational databases in the cloud. This service provides cost-efficient, resizable capacity for an industry-standard relational database and manages common database administration tasks. RDS supports six database engines: Amazon Aurora, PostgreSQL, MySQL, MariaDB, Oracle Database, and SQL Server. These engines give you the ability to run instances ranging from 5GB to 6TB of memory, accommodating your specific use case. It also ensures the database is up-to-date with the latest patches, automatically backs up your data and offers encryption at rest and in transit. +# Glue (ETL) + +AWS Glue is a fully managed ETL service that automates the discovery, cataloging, and transformation of data. It includes a data catalog for storing metadata, a job scheduler, and a serverless Spark environment for running transformations. Glue is commonly used to move and transform data between S3, Redshift, and other AWS services. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/go@4z2i5NXTo9h3YY0kJvRrz.md b/src/data/roadmaps/data-engineer/content/go@4z2i5NXTo9h3YY0kJvRrz.md index 06f4adf47..328d3496a 100644 --- a/src/data/roadmaps/data-engineer/content/go@4z2i5NXTo9h3YY0kJvRrz.md +++ b/src/data/roadmaps/data-engineer/content/go@4z2i5NXTo9h3YY0kJvRrz.md @@ -1,12 +1,10 @@ # Go - -Go, also known as Golang, is a statically typed, compiled programming language designed by Google. It combines the efficiency of compiled languages with the ease of use of dynamically typed interpreted languages. Go features built-in concurrency support through goroutines and channels, making it well-suited for networked and multicore systems. It has a simple and clean syntax, fast compilation times, and efficient garbage collection. Go's standard library is comprehensive, reducing the need for external dependencies. The language emphasizes simplicity and readability, with features like implicit interfaces and a lack of inheritance. Go is particularly popular for building microservices, web servers, and distributed systems. Its performance, simplicity, and robust tooling make it a favored choice for cloud-native development, DevOps tools, and large-scale backend systems. + +Go is a compiled language developed by Google, known for its simplicity, fast execution, and strong concurrency model. In data engineering, it is used to build lightweight, high-throughput services and tools. Its performance characteristics make it a good fit for data pipeline components where latency and resource efficiency matter. Visit the following resources to learn more: - [@roadmap@Visit Dedicated Go Roadmap](https://roadmap.sh/golang) -- [@official@Go Reference Documentation](https://go.dev/doc/) -- [@article@Go by Example - annotated example programs](https://gobyexample.com/) +- [@official@Go Documentation](https://go.dev/doc/) - [@article@Go, the Programming Language of the Cloud](https://thenewstack.io/go-the-programming-language-of-the-cloud/) -- [@video@Go Programming – Golang Course with Bonus Projects](https://www.youtube.com/watch?v=un6ZyFkqFKo) -- [@feed@Explore top posts about Golang](https://app.daily.dev/tags/golang?ref=roadmapsh) \ No newline at end of file +- [@video@Go Programming – Golang Course with Bonus Projects](https://www.youtube.com/watch?v=un6ZyFkqFKo) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/google-cloud-gke@8qEgXYZEbDWC73SQSflDY.md b/src/data/roadmaps/data-engineer/content/google-cloud-gke@8qEgXYZEbDWC73SQSflDY.md index f19b34688..37a678eb2 100644 --- a/src/data/roadmaps/data-engineer/content/google-cloud-gke@8qEgXYZEbDWC73SQSflDY.md +++ b/src/data/roadmaps/data-engineer/content/google-cloud-gke@8qEgXYZEbDWC73SQSflDY.md @@ -1,9 +1,6 @@ -# undefined - -GKE - Google Kubernetes Engine ------------------------------- - -Google Kubernetes Engine (GKE) is a managed Kubernetes service provided by Google Cloud Platform. It allows organizations to deploy, manage, and scale containerized applications using Kubernetes orchestration. GKE automates cluster management tasks, including upgrades, scaling, and security patches, while providing integration with Google Cloud services. It offers features like auto-scaling, load balancing, and private clusters, enabling developers to focus on application development rather than infrastructure management. +# Google Cloud GKE + +Google Kubernetes Engine (GKE) is Google Cloud's managed Kubernetes service. It handles cluster provisioning, upgrades, and scaling automatically, reducing the operational burden of running Kubernetes. GKE is used to run containerized data workloads on Google Cloud infrastructure. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/graph@W6RnhoD7fW2xzVwnyJEDr.md b/src/data/roadmaps/data-engineer/content/graph@W6RnhoD7fW2xzVwnyJEDr.md index 1c2dbd321..190e84423 100644 --- a/src/data/roadmaps/data-engineer/content/graph@W6RnhoD7fW2xzVwnyJEDr.md +++ b/src/data/roadmaps/data-engineer/content/graph@W6RnhoD7fW2xzVwnyJEDr.md @@ -1,12 +1,9 @@ -# Graph Databases - -In a graph database, each node is a record and each arc is a relationship between two nodes. Graph databases are optimized to represent complex relationships with many foreign keys or many-to-many relationships. - -Graphs databases offer high performance for data models with complex relationships, such as a social network. They are relatively new and are not yet widely-used; it might be more difficult to find development tools and resources. Many graphs can only be accessed with REST APIs. +# Graph + +Graph databases store data as nodes and edges, representing entities and the relationships between them. They are optimized for queries that traverse relationships, such as finding connections between users or mapping dependencies. Graph databases are used in social networks, fraud detection, recommendation engines, and knowledge graphs. Visit the following resources to learn more: - [@article@What is a Graph database?](https://aws.amazon.com/nosql/graph/) -- [@article@What is A Graph Database? A Beginner's Guide](https://www.datacamp.com/blog/what-is-a-graph-database) - [@article@Graph database](https://en.wikipedia.org/wiki/Graph_database) - [@video@Introduction to NoSQL](https://www.youtube.com/watch?v=qI_g07C_Q5I) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/hbase@Uho9OOWSG0bUpyH4P6hKk.md b/src/data/roadmaps/data-engineer/content/hbase@Uho9OOWSG0bUpyH4P6hKk.md index 5284799e8..0e116ea33 100644 --- a/src/data/roadmaps/data-engineer/content/hbase@Uho9OOWSG0bUpyH4P6hKk.md +++ b/src/data/roadmaps/data-engineer/content/hbase@Uho9OOWSG0bUpyH4P6hKk.md @@ -1,6 +1,6 @@ # HBase -HBase is a column-oriented No-SQL database management system that runs on top of Hadoop Distributed File System (HDFS), a main component of Apache Hadoop. HBase provides a fault-tolerant way of storing sparse data sets, which are common in many big data use cases. It is well suited for real-time data processing or random read/write access to large volumes of data. HBase applications are written in Java™ much like a typical Apache MapReduce application. +HBase is a column-oriented No-SQL database management system that runs on top of Hadoop Distributed File System (HDFS), a main component of Apache Hadoop. HBase provides a fault-tolerant way of storing sparse data sets, which are common in many big data use cases. It is well-suited for real-time data processing or random read/write access to large volumes of data. HBase applications are written in Java™ much like a typical Apache MapReduce application. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/horizontal-vs-vertical-scaling@k_XSLLwb0Jk0Dd1sw-MpR.md b/src/data/roadmaps/data-engineer/content/horizontal-vs-vertical-scaling@k_XSLLwb0Jk0Dd1sw-MpR.md index 2679bae37..78b76a06e 100644 --- a/src/data/roadmaps/data-engineer/content/horizontal-vs-vertical-scaling@k_XSLLwb0Jk0Dd1sw-MpR.md +++ b/src/data/roadmaps/data-engineer/content/horizontal-vs-vertical-scaling@k_XSLLwb0Jk0Dd1sw-MpR.md @@ -1,8 +1,6 @@ # Horizontal vs Vertical Scaling -Horizontal scaling is the process of adding more machines or nodes to an existing pool in a system to distribute the workload and address increased load. - -By contrast, vertical scaling involves increasing the computing power of individual machines in a system. This is achieved by adjusting or upgrading hardware components, such as CPU, RAM, and network speed. +Horizontal scaling is the process of adding more machines or nodes to an existing pool in a system to distribute the workload and address increased load. By contrast, vertical scaling involves increasing the computing power of individual machines in a system. This is achieved by adjusting or upgrading hardware components, such as CPU, RAM, and network speed. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/indexing@ilbFKqhfYyykjJ7cOngwx.md b/src/data/roadmaps/data-engineer/content/indexing@ilbFKqhfYyykjJ7cOngwx.md index f5a4e6212..c6ce2029a 100644 --- a/src/data/roadmaps/data-engineer/content/indexing@ilbFKqhfYyykjJ7cOngwx.md +++ b/src/data/roadmaps/data-engineer/content/indexing@ilbFKqhfYyykjJ7cOngwx.md @@ -1,5 +1,3 @@ # Indexing - -Indexing is a data structure technique to efficiently retrieve data from a database. It essentially creates a lookup that can be used to quickly find the location of data records on a disk. Indexes are created using a few database columns and are capable of rapidly locating data without scanning every row in a database table each time the database table is accessed. Indexes can be created using any combination of columns in a database table, reducing the amount of time it takes to find data. - -Indexes can be structured in several ways: Binary Tree, B-Tree, Hash Map, etc., each having its own particular strengths and weaknesses. When creating an index, it's crucial to understand which type of index to apply in order to achieve maximum efficiency. Indexes, like any other database feature, must be used wisely because they require disk space and need to be maintained, which can slow down insert and update operations. \ No newline at end of file + +Indexes are data structures that speed up query performance by allowing the database to find rows without scanning the entire table. They are created on one or more columns and come in various types, including B-tree, hash, and full-text indexes. Indexes improve read performance but add overhead to writes and storage. \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/infrastructure-as-code---iac@jgz7L8OSuqRNcf9buuMTj.md b/src/data/roadmaps/data-engineer/content/infrastructure-as-code---iac@jgz7L8OSuqRNcf9buuMTj.md index 83750f0fb..0b9e16c74 100644 --- a/src/data/roadmaps/data-engineer/content/infrastructure-as-code---iac@jgz7L8OSuqRNcf9buuMTj.md +++ b/src/data/roadmaps/data-engineer/content/infrastructure-as-code---iac@jgz7L8OSuqRNcf9buuMTj.md @@ -1,6 +1,6 @@ # Infrastructure as Code - IaC - -Infrastructure as code (IaC) is the ability to provision and support your computing infrastructure using code instead of manual processes and settings. Manual infrastructure management is time-consuming and prone to error—especially when you manage applications at scale. Infrastructure as code lets you define your infrastructure's desired state without including all the steps to get to that state. It automates infrastructure management so developers can focus on building and improving applications instead of managing environments. Organizations use infrastructure as code to control costs, reduce risks, and respond with speed to new business opportunities. + +Infrastructure as Code (IaC) is the practice of managing and provisioning infrastructure through machine-readable configuration files rather than manual processes. Tools like Terraform, AWS CloudFormation, and Pulumi allow teams to define infrastructure declaratively and version it in Git. IaC makes infrastructure reproducible, auditable, and easier to manage at scale. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/integration-testing@NIG53tyoEiLtwf6LvBZId.md b/src/data/roadmaps/data-engineer/content/integration-testing@NIG53tyoEiLtwf6LvBZId.md index 5e90b1ee8..a4b72a3cc 100644 --- a/src/data/roadmaps/data-engineer/content/integration-testing@NIG53tyoEiLtwf6LvBZId.md +++ b/src/data/roadmaps/data-engineer/content/integration-testing@NIG53tyoEiLtwf6LvBZId.md @@ -4,5 +4,4 @@ Integration Testing is a type of testing where software modules are integrated l Visit the following resources to learn more: -- [@article@Integration Testing Tutorial](https://www.guru99.com/integration-testing.html) -- [@feed@Explore top posts about Testing](https://app.daily.dev/tags/testing?ref=roadmapsh) \ No newline at end of file +- [@article@Integration Testing Tutorial](https://www.guru99.com/integration-testing.html) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/introduction@WSYIFni7G2C9Jr0pwuami.md b/src/data/roadmaps/data-engineer/content/introduction@WSYIFni7G2C9Jr0pwuami.md index b3a35f0b5..8f37db337 100644 --- a/src/data/roadmaps/data-engineer/content/introduction@WSYIFni7G2C9Jr0pwuami.md +++ b/src/data/roadmaps/data-engineer/content/introduction@WSYIFni7G2C9Jr0pwuami.md @@ -1,10 +1,7 @@ # Introduction - -Data engineers are responsible for laying the foundations for the acquisition, storage, transformation, and management of data in an organization. They manage the design, creation, and maintenance of database architecture and data processing systems, ensuring that the subsequent work of analysis, BI, and machine learning model development can be carried out seamlessly, continuously, securely, and effectively. - -Data engineers are one of the most technical profiles in the field of data science, bridging the gap between software and application developers and traditional data science positions. + +Data engineering is the discipline of designing, building, and maintaining systems that collect, store, and process data at scale. It sits between raw data sources and the analysts, scientists, and applications that consume that data. The work involves building pipelines, managing storage infrastructure, and ensuring data is reliable and accessible. Visit the following resources to learn more: -- [@article@How to Become a Data Engineer in 2025: 5 Steps for Career Success](https://www.datacamp.com/blog/how-to-become-a-data-engineer) - [@video@What Does a Data Engineer ACTUALLY Do?](https://www.youtube.com/watch?v=hTjo-QVWcK0) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/iot@KeGCHoJRHp-mBX-P5to4Y.md b/src/data/roadmaps/data-engineer/content/iot@KeGCHoJRHp-mBX-P5to4Y.md index 977535614..5679eceb3 100644 --- a/src/data/roadmaps/data-engineer/content/iot@KeGCHoJRHp-mBX-P5to4Y.md +++ b/src/data/roadmaps/data-engineer/content/iot@KeGCHoJRHp-mBX-P5to4Y.md @@ -1,6 +1,6 @@ # IoT -IoT, or Internet of Things, defines a network of connected devices interacting with their environment. IoT devices extend beyond standard devices such as PC's, Laptops or Smartphones, including smart locks, connected thermostats and temperature sensors. In industrial settings, this also includes connected machines, robots, and package tracking devices, and many more. IoT Devices measure and collect data about their environment and some also interact by performing certain predefined actions, for example turning the heat up or down. +IoT, or Internet of Things, refers to a network of connected devices that interact with their environment. IoT devices extend beyond standard devices such as PCs, laptops, and smartphones, including smart locks, connected thermostats, and temperature sensors. In industrial settings, this also includes connected machines, robots, and package tracking devices, and many more. IoT Devices measure and collect data about their environment and some also interact by performing certain predefined actions, for example, turning the heat up or down. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/java@LZ4t8CoCjGWMzE0hScTGZ.md b/src/data/roadmaps/data-engineer/content/java@LZ4t8CoCjGWMzE0hScTGZ.md index 6f9ba9725..c85e1638c 100644 --- a/src/data/roadmaps/data-engineer/content/java@LZ4t8CoCjGWMzE0hScTGZ.md +++ b/src/data/roadmaps/data-engineer/content/java@LZ4t8CoCjGWMzE0hScTGZ.md @@ -4,10 +4,7 @@ Java has had a big influence on data engineering because many core big data tool Visit the following resources to learn more: -- [@book@Thinking in Java](https://www.amazon.co.uk/Thinking-Java-Eckel-Bruce-February/dp/B00IBON6C6) -- [@book@Java: The Complete Reference](https://www.amazon.co.uk/gp/product/B09JL8BMK7/ref=dbs_a_def_rwt_bibl_vppi_i2) -- [@article@@courseIntroduction to Java by Hyperskill (JetBrains Academy)](https://hyperskill.org/courses/8) -- [@article@Effective Java](https://www.amazon.com/Effective-Java-Joshua-Bloch/dp/0134685997) +- [@roadmap@Visit the Dedicated Java Roadmap](https://roadmap.sh/java) +- [@course@Introduction to Java by Hyperskill (JetBrains Academy)](https://hyperskill.org/courses/8) - [@video@Java Tutorial for Beginners](https://www.youtube.com/watch?v=eIrMbAQSU34&feature=youtu.be) -- [@video@Java + DSA + Interview Preparation Course (For beginners)](https://www.youtube.com/playlist?list=PL9gnSGHSqcnr_DxHsP7AW9ftq0AtAyYqJ) -- [@feed@Explore top posts about Java](https://app.daily.dev/tags/java?ref=roadmapsh) \ No newline at end of file +- [@video@Java + DSA + Interview Preparation Course (For beginners)](https://www.youtube.com/playlist?list=PL9gnSGHSqcnr_DxHsP7AW9ftq0AtAyYqJ) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/kubernetes@I_IueX1DFp-LmBwr1-suX.md b/src/data/roadmaps/data-engineer/content/kubernetes@I_IueX1DFp-LmBwr1-suX.md index 270aa76d0..a995e256a 100644 --- a/src/data/roadmaps/data-engineer/content/kubernetes@I_IueX1DFp-LmBwr1-suX.md +++ b/src/data/roadmaps/data-engineer/content/kubernetes@I_IueX1DFp-LmBwr1-suX.md @@ -1,14 +1,10 @@ # Kubernetes - -Kubernetes is an [open source](https://github.com/kubernetes/kubernetes) container management platform, and the dominant product in this space. Using Kubernetes, teams can deploy images across multiple underlying hosts, defining their desired availability, deployment logic, and scaling logic in YAML. Kubernetes evolved from Borg, an internal Google platform used to provision and allocate compute resources (similar to the Autopilot and Aquaman systems of Microsoft Azure). - -The popularity of Kubernetes has made it an increasingly important skill for the DevOps Engineer and has triggered the creation of Platform teams across the industry. These Platform engineering teams often exist with the sole purpose of making Kubernetes approachable and usable for their product development colleagues. + +Kubernetes is an open-source container orchestration system that automates the deployment, scaling, and management of containerized applications. In data engineering, it is used to run pipeline workers, schedule jobs, and manage microservices. Kubernetes has become the standard infrastructure layer for modern data platforms. Visit the following resources to learn more: -- [@official@Kubernetes Website](https://kubernetes.io/) +- [@roadmap@Visit the Dedicated Java Kubernetes](https://roadmap.sh/kubernetes) - [@official@Kubernetes Documentation](https://kubernetes.io/docs/home/) -- [@article@Primer: How Kubernetes Came to Be, What It Is, and Why You Should Care](https://thenewstack.io/primer-how-kubernetes-came-to-be-what-it-is-and-why-you-should-care/) - [@article@Kubernetes: An Overview](https://thenewstack.io/kubernetes-an-overview/) -- [@video@Kubernetes Crash Course for Absolute Beginners](https://www.youtube.com/watch?v=s_o8dwzRlu4) -- [@feed@Explore top posts about Kubernetes](https://app.daily.dev/tags/kubernetes?ref=roadmapsh) \ No newline at end of file +- [@video@Kubernetes Crash Course for Absolute Beginners](https://www.youtube.com/watch?v=s_o8dwzRlu4) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/kubernetes@kcgDW6AFW7WXzXMTPE6J-.md b/src/data/roadmaps/data-engineer/content/kubernetes@kcgDW6AFW7WXzXMTPE6J-.md index 270aa76d0..9df88656d 100644 --- a/src/data/roadmaps/data-engineer/content/kubernetes@kcgDW6AFW7WXzXMTPE6J-.md +++ b/src/data/roadmaps/data-engineer/content/kubernetes@kcgDW6AFW7WXzXMTPE6J-.md @@ -1,14 +1,10 @@ # Kubernetes - -Kubernetes is an [open source](https://github.com/kubernetes/kubernetes) container management platform, and the dominant product in this space. Using Kubernetes, teams can deploy images across multiple underlying hosts, defining their desired availability, deployment logic, and scaling logic in YAML. Kubernetes evolved from Borg, an internal Google platform used to provision and allocate compute resources (similar to the Autopilot and Aquaman systems of Microsoft Azure). - -The popularity of Kubernetes has made it an increasingly important skill for the DevOps Engineer and has triggered the creation of Platform teams across the industry. These Platform engineering teams often exist with the sole purpose of making Kubernetes approachable and usable for their product development colleagues. + +Kubernetes is an open-source system for automating the deployment, scaling, and operation of containerized applications. It manages clusters of containers across multiple nodes and handles load balancing, self-healing, and rolling updates. Data engineers use Kubernetes to run distributed processing jobs, schedule pipeline workers, and manage infrastructure at scale. Visit the following resources to learn more: -- [@official@Kubernetes Website](https://kubernetes.io/) +- [@roadmap@Visit the Dedicated Kubernetes Roadmap](https://roadmap.sh/kubernetes) - [@official@Kubernetes Documentation](https://kubernetes.io/docs/home/) -- [@article@Primer: How Kubernetes Came to Be, What It Is, and Why You Should Care](https://thenewstack.io/primer-how-kubernetes-came-to-be-what-it-is-and-why-you-should-care/) - [@article@Kubernetes: An Overview](https://thenewstack.io/kubernetes-an-overview/) -- [@video@Kubernetes Crash Course for Absolute Beginners](https://www.youtube.com/watch?v=s_o8dwzRlu4) -- [@feed@Explore top posts about Kubernetes](https://app.daily.dev/tags/kubernetes?ref=roadmapsh) \ No newline at end of file +- [@video@Kubernetes Crash Course for Absolute Beginners](https://www.youtube.com/watch?v=s_o8dwzRlu4) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/linux-basics@FXQ_QsljK59zDULLgTqCB.md b/src/data/roadmaps/data-engineer/content/linux-basics@FXQ_QsljK59zDULLgTqCB.md index 00ccae718..dbec5ef30 100644 --- a/src/data/roadmaps/data-engineer/content/linux-basics@FXQ_QsljK59zDULLgTqCB.md +++ b/src/data/roadmaps/data-engineer/content/linux-basics@FXQ_QsljK59zDULLgTqCB.md @@ -7,6 +7,4 @@ Visit the following resources to learn more: - [@roadmap@Visit Dedicated Linux Roadmap](https://roadmap.sh/linux) - [@course@Coursera - Unix Courses](https://www.coursera.org/courses?query=unix) - [@article@Linux Basics](https://dev.to/rudrakshi99/linux-basics-2onj) -- [@article@Unix / Linux Tutorial](https://www.tutorialspoint.com/unix/index.htm) -- [@video@Linux Operating System - Crash Course](https://www.youtube.com/watch?v=ROjZy1WbCIA) -- [@feed@Explore top posts about Linux](https://app.daily.dev/tags/linux?ref=roadmapsh) \ No newline at end of file +- [@video@Linux Operating System - Crash Course](https://www.youtube.com/watch?v=ROjZy1WbCIA) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/load-testing@qoMRpAITA7R_KOrwGDPAb.md b/src/data/roadmaps/data-engineer/content/load-testing@qoMRpAITA7R_KOrwGDPAb.md index ef99abc5f..5bf49941e 100644 --- a/src/data/roadmaps/data-engineer/content/load-testing@qoMRpAITA7R_KOrwGDPAb.md +++ b/src/data/roadmaps/data-engineer/content/load-testing@qoMRpAITA7R_KOrwGDPAb.md @@ -4,5 +4,4 @@ Load Testing is a type of Performance Testing that determines the performance of Visit the following resources to learn more: -- [@article@Load testing and Best Practices](https://loadninja.com/load-testing/) -- [@feed@Explore top posts about Load Testing](https://app.daily.dev/tags/load-testing?ref=roadmapsh) \ No newline at end of file +- [@article@Load testing and Best Practices](https://loadninja.com/load-testing/) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/machine-learning@S8XMtFKWlnUqADElFp0Zw.md b/src/data/roadmaps/data-engineer/content/machine-learning@S8XMtFKWlnUqADElFp0Zw.md index 94cfc0584..a0a42ad8f 100644 --- a/src/data/roadmaps/data-engineer/content/machine-learning@S8XMtFKWlnUqADElFp0Zw.md +++ b/src/data/roadmaps/data-engineer/content/machine-learning@S8XMtFKWlnUqADElFp0Zw.md @@ -4,5 +4,6 @@ Machine learning, a subset of artificial intelligence, is an indispensable tool Visit the following resources to learn more: +- [@roadmap@Visit the Dedicated Java Roadmap](https://roadmap.sh/machine-learning) - [@article@What is Machine Learning (ML)?](https://www.ibm.com/topics/machine-learning) - [@video@What is Machine Learning?](https://www.youtube.com/watch?v=9gGnTQTYNaE) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/mariadb@p7S_6O9Qq722r-F4bl6G3.md b/src/data/roadmaps/data-engineer/content/mariadb@p7S_6O9Qq722r-F4bl6G3.md index 2b2f4e2fe..3e88ac63d 100644 --- a/src/data/roadmaps/data-engineer/content/mariadb@p7S_6O9Qq722r-F4bl6G3.md +++ b/src/data/roadmaps/data-engineer/content/mariadb@p7S_6O9Qq722r-F4bl6G3.md @@ -6,5 +6,4 @@ Visit the following resources to learn more: - [@official@MariaDB](https://mariadb.org/) - [@article@MariaDB vs MySQL](https://www.guru99.com/mariadb-vs-mysql.html) -- [@video@MariaDB Tutorial For Beginners in One Hour](https://www.youtube.com/watch?v=_AMj02sANpI) -- [@feed@Explore top posts about Infrastructure](https://app.daily.dev/tags/infrastructure?ref=roadmapsh) \ No newline at end of file +- [@video@MariaDB Tutorial For Beginners in One Hour](https://www.youtube.com/watch?v=_AMj02sANpI) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/memcached@KYUh29Ok1aeOviboGDS_i.md b/src/data/roadmaps/data-engineer/content/memcached@KYUh29Ok1aeOviboGDS_i.md index eba89bf88..a105e7a9b 100644 --- a/src/data/roadmaps/data-engineer/content/memcached@KYUh29Ok1aeOviboGDS_i.md +++ b/src/data/roadmaps/data-engineer/content/memcached@KYUh29Ok1aeOviboGDS_i.md @@ -1,9 +1,9 @@ # Memcached - -Memcached (pronounced variously mem-cash-dee or mem-cashed) is a general-purpose distributed memory-caching system. It is often used to speed up dynamic database-driven websites by caching data and objects in RAM to reduce the number of times an external data source (such as a database or API) must be read. Memcached is free and open-source software, licensed under the Revised BSD license. Memcached runs on Unix-like operating systems (Linux and macOS) and on Microsoft Windows. It depends on the `libevent` library. Memcached's APIs provide a very large hash table distributed across multiple machines. When the table is full, subsequent inserts cause older data to be purged in the least recently used (LRU) order. Applications using Memcached typically layer requests and additions into RAM before falling back on a slower backing store, such as a database. + +Memcached is a high-performance, distributed in-memory caching system. It is simpler than Redis, supporting only key-value string storage, but is very fast and horizontally scalable. Memcached is commonly used to cache database query results and reduce load on backend systems. Visit the following resources to learn more: -- [@opensource@memcached/memcached](https://github.com/memcached/memcached#readme) +- [@opensource@memcached](https://github.com/memcached/memcached#readme) - [@article@Memcached Tutorial](https://www.tutorialspoint.com/memcached/index.htm) - [@video@Redis vs Memcached](https://www.youtube.com/watch?v=Gyy1SiE8avE) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/messages-vs-streams@IZvL-1Xi0R9IuwJ30FDm4.md b/src/data/roadmaps/data-engineer/content/messages-vs-streams@IZvL-1Xi0R9IuwJ30FDm4.md index 3811d7a68..6fa4ff90b 100644 --- a/src/data/roadmaps/data-engineer/content/messages-vs-streams@IZvL-1Xi0R9IuwJ30FDm4.md +++ b/src/data/roadmaps/data-engineer/content/messages-vs-streams@IZvL-1Xi0R9IuwJ30FDm4.md @@ -1,5 +1,3 @@ # Messages vs Streams - -Messages and Streams are often used interchange‐ably but a subtle but essential differences exists between the two. A message is raw data communicated across two or more systems. Messages are discrete and singular signals in an event-driven system. - -By contrast, a stream is an append-only log of event records. As events occur, streams are accumulated in an ordered sequence, using a timestamp or an ID to record events order. Streams are used when you need to analyze what happened over many events. Because of the append-only nature of streams, records in a stream are persisted over a long retention window—often weeks or months—allowing for complex operations on records such as aggregations on multiple records or the ability to rewind to a point in time within the stream. \ No newline at end of file + +Messages are discrete, individual units of data sent from a producer to a consumer. Streams are continuous, ordered sequences of data that consumers process in real time or replay from a position. Systems like RabbitMQ focus on message delivery, while Apache Kafka is designed around the stream abstraction and supports log retention and replay. \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/metadata-first-architecture@14CycunRC1p2qTRn-ncoy.md b/src/data/roadmaps/data-engineer/content/metadata-first-architecture@14CycunRC1p2qTRn-ncoy.md index 96f6126bd..820b76419 100644 --- a/src/data/roadmaps/data-engineer/content/metadata-first-architecture@14CycunRC1p2qTRn-ncoy.md +++ b/src/data/roadmaps/data-engineer/content/metadata-first-architecture@14CycunRC1p2qTRn-ncoy.md @@ -1 +1,3 @@ -# Metadata-first Architecture \ No newline at end of file +# Metadata-first Architecture + +A metadata-first architecture treats metadata as a first-class citizen in data system design. Rather than just documenting data after the fact, metadata is captured and used actively to govern, discover, and lineage-track data across the organization. This approach supports better data quality, compliance, and self-service analytics. \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/metadata-management@a5gzM8msXibxD58eVDkM-.md b/src/data/roadmaps/data-engineer/content/metadata-management@a5gzM8msXibxD58eVDkM-.md index 1b7dcea32..2ebbdc5ad 100644 --- a/src/data/roadmaps/data-engineer/content/metadata-management@a5gzM8msXibxD58eVDkM-.md +++ b/src/data/roadmaps/data-engineer/content/metadata-management@a5gzM8msXibxD58eVDkM-.md @@ -1 +1,3 @@ -# Metadata Management \ No newline at end of file +# Metadata Management + +Metadata management involves capturing, storing, and governing information about data assets, including their structure, origin, ownership, and usage. A metadata catalog makes data discoverable and helps teams understand what data is available and how it is defined. Tools like Apache Atlas, DataHub, and Alation are used for metadata management. \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/microsoft-power-bi@6Nr5FAGT_oOPZwZWdv7hl.md b/src/data/roadmaps/data-engineer/content/microsoft-power-bi@6Nr5FAGT_oOPZwZWdv7hl.md index 0a9214010..d0382c67c 100644 --- a/src/data/roadmaps/data-engineer/content/microsoft-power-bi@6Nr5FAGT_oOPZwZWdv7hl.md +++ b/src/data/roadmaps/data-engineer/content/microsoft-power-bi@6Nr5FAGT_oOPZwZWdv7hl.md @@ -1,6 +1,6 @@ # Microsoft Power BI - -PowerBI, an interactive data visualization and business analytics tool developed by Microsoft, plays a crucial role in the field of a data analyst's work. It helps data analysts to convert raw data into meaningful insights through it's easy-to-use dashboards and reports function. This tool provides a unified view of business data, allowing analysts to track and visualize key performance metrics and make better-informed business decisions. With PowerBI, data analysts also have the ability to manipulate and produce visualizations of large data sets that can be shared across an organization, making complex statistical information more digestible. + +Microsoft Power BI is a business intelligence and data visualization platform from Microsoft. It connects to a wide range of data sources and allows users to build interactive dashboards and reports. Power BI integrates tightly with the Microsoft ecosystem including Azure and Excel. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/mlops@VQv-c7buU2l-IDzRZBMRo.md b/src/data/roadmaps/data-engineer/content/mlops@VQv-c7buU2l-IDzRZBMRo.md index c7e0a2c74..8c834441b 100644 --- a/src/data/roadmaps/data-engineer/content/mlops@VQv-c7buU2l-IDzRZBMRo.md +++ b/src/data/roadmaps/data-engineer/content/mlops@VQv-c7buU2l-IDzRZBMRo.md @@ -1,3 +1,7 @@ # MLOps -MLOps is a practice for collaboration and communication between data scientists and operations professionals to help manage production ML lifecycle. It is a set of best practices that aims to automate the ML lifecycle, including training, deployment, and monitoring. MLOps helps organizations to scale ML models and deliver business value faster. \ No newline at end of file +MLOps is a practice for collaboration and communication between data scientists and operations professionals to help manage production ML lifecycle. It is a set of best practices that aims to automate the ML lifecycle, including training, deployment, and monitoring. MLOps helps organizations to scale ML models and deliver business value faster. + +Visit the following resources to learn more: + +- [@roadmap@Visit the Dedicated MLOps Roadmap](https://roadmap.sh/mlops) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/mongodb@04V0Bcgjusfqdw0b-Aw4W.md b/src/data/roadmaps/data-engineer/content/mongodb@04V0Bcgjusfqdw0b-Aw4W.md index d85687df8..45407fd66 100644 --- a/src/data/roadmaps/data-engineer/content/mongodb@04V0Bcgjusfqdw0b-Aw4W.md +++ b/src/data/roadmaps/data-engineer/content/mongodb@04V0Bcgjusfqdw0b-Aw4W.md @@ -7,5 +7,4 @@ Visit the following resources to learn more: - [@roadmap@Visit Dedicated MongoDB Roadmap](https://roadmap.sh/mongodb) - [@official@MongoDB Website](https://www.mongodb.com/) - [@official@Learning Path for MongoDB Developers](https://learn.mongodb.com/catalog) -- [@article@MongoDB Online Sandbox](https://mongoplayground.net/) -- [@feed@daily.dev MongoDB Feed](https://app.daily.dev/tags/mongodb) \ No newline at end of file +- [@article@MongoDB Online Sandbox](https://mongoplayground.net/) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/monitoring@dk5FQl7Pk3-O5eF7dKwmp.md b/src/data/roadmaps/data-engineer/content/monitoring@dk5FQl7Pk3-O5eF7dKwmp.md index 30b88fcf4..ae13e3a00 100644 --- a/src/data/roadmaps/data-engineer/content/monitoring@dk5FQl7Pk3-O5eF7dKwmp.md +++ b/src/data/roadmaps/data-engineer/content/monitoring@dk5FQl7Pk3-O5eF7dKwmp.md @@ -1,10 +1,7 @@ # Monitoring - -Monitoring involves continuously observing and tracking the performance, availability, and health of systems, applications, and infrastructure. It typically includes collecting and analyzing metrics, logs, and events to ensure systems are operating within desired parameters. Monitoring helps detect anomalies, identify potential issues before they escalate, and provides insights into system behavior. It often involves tools and platforms that offer dashboards, alerts, and reporting features to facilitate real-time visibility and proactive management. Effective monitoring is crucial for maintaining system reliability, performance, and for supporting incident response and troubleshooting. - -A few popular tools are Prometheus, Sentry, Datadog, and NewRelic. + +Monitoring is the practice of collecting and analyzing metrics, logs, and events from running systems to track their health and performance. For data pipelines, monitoring covers job success rates, data freshness, latency, and resource usage. Good monitoring enables fast detection and diagnosis of failures before they affect downstream consumers. Visit the following resources to learn more: -- [@article@Top Monitoring Tools](https://thectoclub.com/tools/best-application-monitoring-software/) -- [@feed@daily.dev Monitoring Feed](https://app.daily.dev/tags/monitoring) \ No newline at end of file +- [@article@Top Monitoring Tools](https://thectoclub.com/tools/best-application-monitoring-software/) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/ms-sql@YxnIQh6Y5ic795-YsajB8.md b/src/data/roadmaps/data-engineer/content/ms-sql@YxnIQh6Y5ic795-YsajB8.md index 9111ae72d..8465a0089 100644 --- a/src/data/roadmaps/data-engineer/content/ms-sql@YxnIQh6Y5ic795-YsajB8.md +++ b/src/data/roadmaps/data-engineer/content/ms-sql@YxnIQh6Y5ic795-YsajB8.md @@ -1,6 +1,6 @@ # MS SQL - -Microsoft SQL Server (MS SQL) is a relational database management system developed by Microsoft for managing and storing structured data. It supports a wide range of data operations, including querying, transaction management, and data warehousing. SQL Server provides tools and features for database design, performance optimization, and security, including support for complex queries through T-SQL (Transact-SQL), data integration with SQL Server Integration Services (SSIS), and business intelligence with SQL Server Analysis Services (SSAS) and SQL Server Reporting Services (SSRS). It is commonly used in enterprise environments for applications requiring reliable data storage, transaction processing, and reporting. + +Microsoft SQL Server (MS SQL) is a relational database developed by Microsoft, commonly used in enterprise and Windows-based environments. It integrates tightly with the Microsoft ecosystem, including Azure, Power BI, and .NET. MS SQL supports T-SQL, Microsoft's extension of SQL with additional procedural capabilities. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/mysql@_bFj6rbLuqeQB5MjJZpd6.md b/src/data/roadmaps/data-engineer/content/mysql@_bFj6rbLuqeQB5MjJZpd6.md index 7e2c581dc..936b56c60 100644 --- a/src/data/roadmaps/data-engineer/content/mysql@_bFj6rbLuqeQB5MjJZpd6.md +++ b/src/data/roadmaps/data-engineer/content/mysql@_bFj6rbLuqeQB5MjJZpd6.md @@ -7,5 +7,4 @@ Visit the following resources to learn more: - [@official@MySQL](https://www.mysql.com/) - [@article@MySQL for Developers](https://planetscale.com/courses/mysql-for-developers/introduction/course-introduction) - [@article@MySQL Tutorial](https://www.mysqltutorial.org/) -- [@video@MySQL Complete Course](https://www.youtube.com/watch?v=5OdVJbNCSso) -- [@feed@Explore top posts about MySQL](https://app.daily.dev/tags/mysql?ref=roadmapsh) \ No newline at end of file +- [@video@MySQL Complete Course](https://www.youtube.com/watch?v=5OdVJbNCSso) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/neo4j@TG63YRbSKL1F9vlUVF1VY.md b/src/data/roadmaps/data-engineer/content/neo4j@TG63YRbSKL1F9vlUVF1VY.md index 50af36728..c8e75d688 100644 --- a/src/data/roadmaps/data-engineer/content/neo4j@TG63YRbSKL1F9vlUVF1VY.md +++ b/src/data/roadmaps/data-engineer/content/neo4j@TG63YRbSKL1F9vlUVF1VY.md @@ -1,10 +1,9 @@ -# NEO4J - -Neo4j is a highly popular open-source graph database designed to store, manage, and query data as interconnected nodes and relationships. Unlike traditional relational databases that use tables and rows, Neo4j uses a graph model where data is represented as nodes (entities) and edges (relationships), allowing for highly efficient querying of complex, interconnected data. It supports Cypher, a declarative query language specifically designed for graph querying, which simplifies operations like traversing relationships and pattern matching. Neo4j is well-suited for applications involving complex relationships, such as social networks, recommendation engines, and fraud detection, where understanding and leveraging connections between data points is crucial. +# Neo4j + +Neo4j is the most widely used graph database. It stores data natively as nodes and relationships and uses Cypher, a declarative query language designed for graph traversal. Neo4j is used for applications where relationship-heavy queries are central, such as recommendation systems and network analysis. Visit the following resources to learn more: - [@official@Neo4j Website](https://neo4j.com) - [@video@Neo4j in 100 Seconds](https://www.youtube.com/watch?v=T6L9EoBy8Zk) -- [@video@Neo4j Course for Beginners](https://www.youtube.com/watch?v=_IgbB24scLI) -- [@feed@Explore top posts about Backend Development](https://app.daily.dev/tags/backend?ref=roadmapsh) \ No newline at end of file +- [@video@Neo4j Course for Beginners](https://www.youtube.com/watch?v=_IgbB24scLI) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/neptune@atAK4zGXIbxZvfBTzFEIe.md b/src/data/roadmaps/data-engineer/content/neptune@atAK4zGXIbxZvfBTzFEIe.md index 12ae67ebb..2943bf086 100644 --- a/src/data/roadmaps/data-engineer/content/neptune@atAK4zGXIbxZvfBTzFEIe.md +++ b/src/data/roadmaps/data-engineer/content/neptune@atAK4zGXIbxZvfBTzFEIe.md @@ -1,6 +1,6 @@ -# AWS Neptune - -Amazon Neptune is a fully managed graph database service provided by Amazon Web Services (AWS). It's designed to store and navigate highly connected data, supporting both property graph and RDF (Resource Description Framework) models. Neptune uses graph query languages like Gremlin and SPARQL, making it suitable for applications involving complex relationships, such as social networks, recommendation engines, fraud detection systems, and knowledge graphs. It offers high availability, with replication across multiple Availability Zones, and supports up to 15 read replicas for improved performance. Neptune integrates with other AWS services, provides encryption at rest and in transit, and offers fast recovery from failures. Its scalability and performance make it valuable for handling large-scale, complex data relationships in enterprise-level applications. +# Neptune + +Amazon Neptune is a managed graph database service on AWS that supports both the Property Graph model (with Gremlin) and RDF (with SPARQL). It is designed for highly connected datasets and scales to billions of relationships. Neptune is used for knowledge graphs, fraud detection, and identity resolution. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/networking-fundamentals@cgkzFMmQils2sYj4NW8VW.md b/src/data/roadmaps/data-engineer/content/networking-fundamentals@cgkzFMmQils2sYj4NW8VW.md index c78249c32..ed2a3c82e 100644 --- a/src/data/roadmaps/data-engineer/content/networking-fundamentals@cgkzFMmQils2sYj4NW8VW.md +++ b/src/data/roadmaps/data-engineer/content/networking-fundamentals@cgkzFMmQils2sYj4NW8VW.md @@ -1,12 +1,10 @@ -# Networking - -Networking is the process of connecting two or more computing devices together for the purpose of sharing data. In a data network, shared data may be as simple as a printer or as complex as a global financial transaction. - -If you have networking experience or want to be a reliability engineer or operations engineer, expect questions from these topics. Otherwise, this is just good to know. +Networking Fundamentals + +Networking fundamentals cover how data moves between systems: IP addressing, DNS, HTTP, TCP/UDP, and firewalls. Data engineers encounter networking when configuring cloud resources, troubleshooting pipeline failures, or setting up secure connections between services. A basic understanding of how networks operate helps diagnose connectivity issues and design reliable architectures. Visit the following resources to learn more: +- [@roadmap@Visit the Dedicated Network Engineer Roadmap](https://roadmap.sh/network-engineer) - [@article@Khan Academy - Networking](https://www.khanacademy.org/computing/code-org/computers-and-the-internet) - [@video@Computer Networking Course - Network Engineering](https://www.youtube.com/watch?v=qiQR5rTSshw) -- [@video@Networking Video Series (21 videos)](https://www.youtube.com/playlist?list=PLEbnTDJUr_IegfoqO4iPnPYQui46QqT0j) -- [@feed@Explore top posts about Networking](https://app.daily.dev/tags/networking?ref=roadmapsh) \ No newline at end of file +- [@video@Networking Video Series (21 videos)](https://www.youtube.com/playlist?list=PLEbnTDJUr_IegfoqO4iPnPYQui46QqT0j) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/new-relic@r1KmASWAa_MOqQOC9gvvF.md b/src/data/roadmaps/data-engineer/content/new-relic@r1KmASWAa_MOqQOC9gvvF.md index 51d677e8c..08f01617c 100644 --- a/src/data/roadmaps/data-engineer/content/new-relic@r1KmASWAa_MOqQOC9gvvF.md +++ b/src/data/roadmaps/data-engineer/content/new-relic@r1KmASWAa_MOqQOC9gvvF.md @@ -4,6 +4,4 @@ New Relic is an observability platform that helps you build better software. You Visit the following resources to learn more: -- [@official@New Relic](https://newrelic.com/) -- [@official@Learn New Relic](https://learn.newrelic.com/) -- [@feed@Explore top posts about DevOps](https://app.daily.dev/tags/devops?ref=roadmapsh) \ No newline at end of file +- [@official@Learn New Relic](https://learn.newrelic.com/) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/nosql-databases@uZYQ8tqTriXt_JIOjcM9_.md b/src/data/roadmaps/data-engineer/content/nosql-databases@uZYQ8tqTriXt_JIOjcM9_.md new file mode 100644 index 000000000..c739c5120 --- /dev/null +++ b/src/data/roadmaps/data-engineer/content/nosql-databases@uZYQ8tqTriXt_JIOjcM9_.md @@ -0,0 +1,9 @@ +# NoSQL Databases + +NoSQL databases are a category of database systems that do not use the traditional relational table model. They are designed for specific data models and access patterns, including document storage, key-value pairs, wide-column stores, and graphs. NoSQL databases are often chosen for their scalability, flexibility, and performance in high-volume or unstructured data scenarios. + +Visit the following resources to learn more: + +- [@article@NoSQL Explained](https://www.mongodb.com/nosql-explained) +- [@video@How do NoSQL Databases work](https://www.youtube.com/watch?v=0buKQHokLK8) +- [@video@SQL vs NoSQL Explained](https://www.youtube.com/watch?v=ruz-vK8IesE) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/oltp-vs-olap@-VQQmIUGesnrT1N6kH5et.md b/src/data/roadmaps/data-engineer/content/oltp-vs-olap@-VQQmIUGesnrT1N6kH5et.md index b6113fd44..33e6a9d92 100644 --- a/src/data/roadmaps/data-engineer/content/oltp-vs-olap@-VQQmIUGesnrT1N6kH5et.md +++ b/src/data/roadmaps/data-engineer/content/oltp-vs-olap@-VQQmIUGesnrT1N6kH5et.md @@ -1,8 +1,6 @@ # OLTP vs OLAP - -Online Transaction Processing (OLTP) refers to a class of systems designed to manage transaction-oriented applications, typically for data entry and retrieval transactions in database systems. OLTP systems are characterized by a large number of short online transactions (INSERT, UPDATE, DELETE), where the emphasis is on speed, efficiency, and maintaining data integrity in multi-access environments. PostgreSQL supports OLTP workloads through features like ACID compliance (Atomicity, Consistency, Isolation, Durability), MVCC (Multi-Version Concurrency Control) for high concurrency, efficient indexing, and robust transaction management. These features ensure reliable, fast, and consistent processing of high-volume, high-frequency transactions critical to OLTP applications. - -Online Analytical Processing (OLAP) refers to a class of systems designed for query-intensive tasks, typically used for data analysis and business intelligence. OLAP systems handle complex queries that aggregate large volumes of data, often from multiple sources, to support decision-making processes. + +OLTP (Online Transaction Processing) systems are optimized for fast, frequent read and write operations, typically backing operational applications. OLAP (Online Analytical Processing) systems are designed for complex queries over large datasets, used in reporting and analytics. The two have different storage formats, indexing strategies, and performance characteristics. Visit the following resources to learn more: diff --git a/src/data/roadmaps/data-engineer/content/oracle@PJcxM60h85Po0AAkSj7nr.md b/src/data/roadmaps/data-engineer/content/oracle@PJcxM60h85Po0AAkSj7nr.md index c2c404997..b19077d38 100644 --- a/src/data/roadmaps/data-engineer/content/oracle@PJcxM60h85Po0AAkSj7nr.md +++ b/src/data/roadmaps/data-engineer/content/oracle@PJcxM60h85Po0AAkSj7nr.md @@ -1,10 +1,8 @@ # Oracle - -Oracle Database is a highly robust, enterprise-grade relational database management system (RDBMS) developed by Oracle Corporation. Known for its scalability, reliability, and comprehensive features, Oracle Database supports complex data management tasks and mission-critical applications. It provides advanced functionalities like SQL querying, transaction management, high availability through clustering, and data warehousing. Oracle's database solutions include support for various data models, such as relational, spatial, and graph, and offer tools for security, performance optimization, and data integration. It is widely used in industries requiring large-scale, secure, and high-performance data processing. + +Oracle Database is a commercial relational database system widely used in enterprise environments. It is known for its robustness, advanced features, and support for very large-scale deployments. Oracle is common in financial services, healthcare, and government sectors where long-term vendor support and mature tooling are priorities. Visit the following resources to learn more: -- [@official@Oracle Website](https://www.oracle.com/database/) - [@official@Oracle Docs](https://docs.oracle.com/en/database/index.html) -- [@video@Oracle SQL Tutorial for Beginners](https://www.youtube.com/watch?v=ObbNGhcxXJA) -- [@feed@Explore top posts about Oracle](https://app.daily.dev/tags/oracle?ref=roadmapsh) \ No newline at end of file +- [@video@Oracle SQL Tutorial for Beginners](https://www.youtube.com/watch?v=ObbNGhcxXJA) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/postgresql@__JFgwxeDLvz8p7DAJnsc.md b/src/data/roadmaps/data-engineer/content/postgresql@__JFgwxeDLvz8p7DAJnsc.md index 71aba314f..1bd8cbcff 100644 --- a/src/data/roadmaps/data-engineer/content/postgresql@__JFgwxeDLvz8p7DAJnsc.md +++ b/src/data/roadmaps/data-engineer/content/postgresql@__JFgwxeDLvz8p7DAJnsc.md @@ -1,12 +1,10 @@ # PostgreSQL - -PostgreSQL is an advanced, open-source relational database management system (RDBMS) known for its robustness, extensibility, and standards compliance. It supports a wide range of data types and advanced features, including complex queries, foreign keys, and full-text search. PostgreSQL is highly extensible, allowing users to define custom data types, operators, and functions. It supports ACID (Atomicity, Consistency, Isolation, Durability) properties for reliable transaction processing and offers strong support for concurrency and data integrity. Its capabilities make it suitable for various applications, from simple web apps to large-scale data warehousing and analytics solutions. + +PostgreSQL is an open-source relational database known for its standards compliance, extensibility, and advanced feature set. It supports complex queries, JSON storage, full-text search, and custom data types. PostgreSQL is widely used in both transactional and analytical workloads. Visit the following resources to learn more: - [@roadmap@Visit Dedicated PostgreSQL DBA Roadmap](https://roadmap.sh/postgresql-dba) -- [@official@Official Website](https://www.postgresql.org/) +- [@official@PostgreSQL Website](https://www.postgresql.org/) - [@article@Learn PostgreSQL - Full Tutorial for Beginners](https://www.postgresqltutorial.com/) -- [@video@PostgreSQL in 100 Seconds](https://www.youtube.com/watch?v=n2Fluyr3lbc) -- [@video@Postgres tutorial for Beginners](https://www.youtube.com/watch?v=SpfIwlAYaKk) -- [@feed@Explore top posts about PostgreSQL](https://app.daily.dev/tags/postgresql?ref=roadmapsh) \ No newline at end of file +- [@video@Postgres tutorial for Beginners](https://www.youtube.com/watch?v=SpfIwlAYaKk) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/prefect@TAh4__7U58J7fduU9a1Ol.md b/src/data/roadmaps/data-engineer/content/prefect@TAh4__7U58J7fduU9a1Ol.md new file mode 100644 index 000000000..1bc930c9d --- /dev/null +++ b/src/data/roadmaps/data-engineer/content/prefect@TAh4__7U58J7fduU9a1Ol.md @@ -0,0 +1,8 @@ +# Prefect + +Prefect is an open-source orchestration engine that turns your Python functions into production-grade data pipelines with minimal friction. You can build and schedule workflows in pure Python—no DSLs or complex config files—and run them anywhere you can run Python. Prefect handles the heavy lifting for you out of the box: automatic state tracking, failure handling, real-time monitoring, and more. + +Visit the following resources to learn more: + +- [@official@Prefect Docs](https://docs.prefect.io/v3/get-started) +- [@video@Getting Started with Prefect](https://www.youtube.com/watch?v=D5DhwVNHWeU) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/programming-skills@_2Ofq3Df-VRXDgKyveZ0U.md b/src/data/roadmaps/data-engineer/content/programming-skills@_2Ofq3Df-VRXDgKyveZ0U.md index 3ba2bc633..da04dea53 100644 --- a/src/data/roadmaps/data-engineer/content/programming-skills@_2Ofq3Df-VRXDgKyveZ0U.md +++ b/src/data/roadmaps/data-engineer/content/programming-skills@_2Ofq3Df-VRXDgKyveZ0U.md @@ -1,3 +1,3 @@ # Programming Skills -To be successful as a data engineer, you need to be proficient in coding. This involves knowning basic concepts and principles that form the foundation of any computer programming language. These include understanding variables, which store data for processing, control structures such as loops and conditional statements that direct the flow of a program, data structures which organize and store data efficiently, and algorithms which step by step instructions to solve specific problems or perform specific tasks. \ No newline at end of file +To be successful as a data engineer, you need to be proficient in coding. This involves knowing basic concepts and principles that form the foundation of any computer programming language. These include understanding variables, which store data for processing, control structures such as loops and conditional statements that direct the flow of a program, data structures which organize and store data efficiently, and algorithms which provide step-by-step instructions to solve specific problems or perform specific tasks. \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/prometheus@3QsgoKKxAoyj2LWJ8ad-7.md b/src/data/roadmaps/data-engineer/content/prometheus@3QsgoKKxAoyj2LWJ8ad-7.md index 33d8e6aa5..63c8bed73 100644 --- a/src/data/roadmaps/data-engineer/content/prometheus@3QsgoKKxAoyj2LWJ8ad-7.md +++ b/src/data/roadmaps/data-engineer/content/prometheus@3QsgoKKxAoyj2LWJ8ad-7.md @@ -4,7 +4,5 @@ Prometheus is a free software application used for event monitoring and alerting Visit the following resources to learn more: -- [@official@Prometheus Website](https://prometheus.io/) - [@official@Prometheus Documentation](https://prometheus.io/docs/introduction/overview/) -- [@official@Getting Started with Prometheus](https://prometheus.io/docs/tutorials/getting_started/) -- [@feed@Explore top posts about Prometheus](https://app.daily.dev/tags/prometheus?ref=roadmapsh) \ No newline at end of file +- [@official@Getting Started with Prometheus](https://prometheus.io/docs/tutorials/getting_started/) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/python@ILs5azr4L_uLK0CDFKVaz.md b/src/data/roadmaps/data-engineer/content/python@ILs5azr4L_uLK0CDFKVaz.md index 1a2e9bdfa..6e4edc1cf 100644 --- a/src/data/roadmaps/data-engineer/content/python@ILs5azr4L_uLK0CDFKVaz.md +++ b/src/data/roadmaps/data-engineer/content/python@ILs5azr4L_uLK0CDFKVaz.md @@ -4,9 +4,7 @@ Python’s inherent characteristics and the wealth of resources that have grown Visit the following resources to learn more: -- [@official@Python Website](https://www.python.org/) -- [@article@Python - Wiki](https://en.wikipedia.org/wiki/Python_(programming_language)) +- [@roadmap@Visit the Dedicated Python Roadmap](https://roadmap.sh/python) - [@article@Tutorial Series: How to Code in Python](https://www.digitalocean.com/community/tutorials/how-to-write-your-first-python-3-program) - [@article@Google's Python Class](https://developers.google.com/edu/python) -- [@video@Learn Python - Full Course](https://www.youtube.com/watch?v=4M87qBgpafk) -- [@feed@Explore top posts about Python](https://app.daily.dev/tags/python?ref=roadmapsh) \ No newline at end of file +- [@video@Learn Python - Full Course](https://www.youtube.com/watch?v=4M87qBgpafk) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/rabbitmq@ERcgPTACqYo9BXoRdLjbd.md b/src/data/roadmaps/data-engineer/content/rabbitmq@ERcgPTACqYo9BXoRdLjbd.md index c33d56a11..2494c750f 100644 --- a/src/data/roadmaps/data-engineer/content/rabbitmq@ERcgPTACqYo9BXoRdLjbd.md +++ b/src/data/roadmaps/data-engineer/content/rabbitmq@ERcgPTACqYo9BXoRdLjbd.md @@ -1,10 +1,9 @@ # RabbitMQ - -RabbitMQ is an open-source message broker that facilitates the exchange of messages between distributed systems using the Advanced Message Queuing Protocol (AMQP). It enables asynchronous communication by queuing and routing messages between producers and consumers, which helps decouple application components and improve scalability and reliability. RabbitMQ supports features such as message durability, acknowledgments, and flexible routing through exchanges and queues. It is highly configurable, allowing for various messaging patterns, including publish/subscribe, request/reply, and point-to-point communication. RabbitMQ is widely used in enterprise environments for handling high-throughput messaging and integrating heterogeneous systems. + +RabbitMQ is an open-source message broker that implements the AMQP protocol. It routes messages between producers and consumers using exchanges and queues, supporting patterns like publish/subscribe, work queues, and routing. RabbitMQ is used for task queues, service-to-service communication, and event notification systems. Visit the following resources to learn more: - [@official@RabbitMQ Tutorials](https://www.rabbitmq.com/getstarted.html) - [@video@RabbitMQ Tutorial - Message Queues and Distributed Systems](https://www.youtube.com/watch?v=nFxjaVmFj5E) -- [@video@RabbitMQ in 100 Seconds](https://m.youtube.com/watch?v=NQ3fZtyXji0) -- [@feed@Explore top posts about RabbitMQ](https://app.daily.dev/tags/rabbitmq?ref=roadmapsh) \ No newline at end of file +- [@video@RabbitMQ in 100 Seconds](https://m.youtube.com/watch?v=NQ3fZtyXji0) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/redis@dW_eC4vR8BrvKG9wxmEBc.md b/src/data/roadmaps/data-engineer/content/redis@dW_eC4vR8BrvKG9wxmEBc.md index ecfe6aa44..672fbe3eb 100644 --- a/src/data/roadmaps/data-engineer/content/redis@dW_eC4vR8BrvKG9wxmEBc.md +++ b/src/data/roadmaps/data-engineer/content/redis@dW_eC4vR8BrvKG9wxmEBc.md @@ -1,12 +1,10 @@ # Redis - -Redis is an open-source, in-memory data structure store known for its speed and versatility. It supports various data types, including strings, lists, sets, hashes, and sorted sets, and provides functionalities such as caching, session management, real-time analytics, and message brokering. Redis operates as a key-value store, allowing for rapid read and write operations, and is often used to enhance performance and scalability in applications. It supports persistence options to save data to disk, replication for high availability, and clustering for horizontal scaling. Redis is widely used for scenarios requiring low-latency access to data and high-throughput performance. + +Redis is an in-memory key-value store known for its extremely low latency. It supports a variety of data structures including strings, lists, sets, sorted sets, and hashes. Redis is widely used for caching, real-time leaderboards, pub/sub messaging, and session storage. Visit the following resources to learn more: - [@roadmap@Visit Dedicated Redis Roadmap](https://roadmap.sh/redis) - [@course@Redis Crash Course](https://www.youtube.com/watch?v=XCsS_NVAa1g) -- [@official@Redis](https://redis.io/) - [@official@Redis Documentation](https://redis.io/docs/latest/) -- [@video@Redis in 100 Seconds](https://www.youtube.com/watch?v=G1rOthIU-uo) -- [@feed@Explore top posts about Redis](https://app.daily.dev/tags/redis?ref=roadmapsh) \ No newline at end of file +- [@video@Redis in 100 Seconds](https://www.youtube.com/watch?v=G1rOthIU-uo) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/relational-databases@cslVSSKBMO7I6CpO7vG1H.md b/src/data/roadmaps/data-engineer/content/relational-databases@cslVSSKBMO7I6CpO7vG1H.md index f6f3a629b..c7d60a66d 100644 --- a/src/data/roadmaps/data-engineer/content/relational-databases@cslVSSKBMO7I6CpO7vG1H.md +++ b/src/data/roadmaps/data-engineer/content/relational-databases@cslVSSKBMO7I6CpO7vG1H.md @@ -1,12 +1,10 @@ # Relational Databases - -Relational databases are a type of database management system (DBMS) that organizes data into structured tables with rows and columns, using a schema to define data relationships and constraints. They employ Structured Query Language (SQL) for querying and managing data, supporting operations such as data retrieval, insertion, updating, and deletion. Relational databases enforce data integrity through keys (primary and foreign) and constraints (such as unique and not-null), and they are designed to handle complex queries, transactions, and data relationships efficiently. Examples of relational databases include MySQL, PostgreSQL, and Oracle Database. They are commonly used for applications requiring structured data storage, strong consistency, and complex querying capabilities. + +Relational databases store data in tables with rows and columns, and use SQL for querying. Relationships between tables are defined through foreign keys. They are the most widely used type of database for transactional applications and form the backbone of most operational systems. Visit the following resources to learn more: - [@course@Databases and SQL](https://www.edx.org/course/databases-5-sql) - [@article@Relational Databases](https://www.ibm.com/cloud/learn/relational-databases) -- [@article@51 Years of Relational Databases](https://learnsql.com/blog/codd-article-databases/) - [@article@Intro To Relational Databases](https://www.udacity.com/course/intro-to-relational-databases--ud197) -- [@video@What is Relational Database](https://youtu.be/OqjJjpjDRLc) -- [@feed@Explore top posts about Backend Development](https://app.daily.dev/tags/backend?ref=roadmapsh) \ No newline at end of file +- [@video@What is Relational Database](https://youtu.be/OqjJjpjDRLc) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/reverse-etl-usecases@mBOGrJIUaatBe2PnJM2NK.md b/src/data/roadmaps/data-engineer/content/reverse-etl-usecases@mBOGrJIUaatBe2PnJM2NK.md index 2fb412203..c38f7776c 100644 --- a/src/data/roadmaps/data-engineer/content/reverse-etl-usecases@mBOGrJIUaatBe2PnJM2NK.md +++ b/src/data/roadmaps/data-engineer/content/reverse-etl-usecases@mBOGrJIUaatBe2PnJM2NK.md @@ -1 +1,3 @@ -# Reverse ETL Usecases \ No newline at end of file +# Reverse ETL Usecases + +Common use cases for reverse ETL include syncing customer health scores to a CRM, pushing segmented user lists to a marketing automation platform, and sending product usage data to customer success tools. It enables business teams to act on insights derived in the data warehouse without needing access to it directly. \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/reverse-etl@JpuiYsipNWBcrjmn2ji6b.md b/src/data/roadmaps/data-engineer/content/reverse-etl@JpuiYsipNWBcrjmn2ji6b.md index 55652679c..08e980f14 100644 --- a/src/data/roadmaps/data-engineer/content/reverse-etl@JpuiYsipNWBcrjmn2ji6b.md +++ b/src/data/roadmaps/data-engineer/content/reverse-etl@JpuiYsipNWBcrjmn2ji6b.md @@ -4,5 +4,4 @@ Reverse ETL is the process of extracting data from a data warehouse, transformin Visit the following resources to learn more: -- [@article@What is Reverse ETL? A Helpful Guide](https://www.datacamp.com/blog/reverse-etl) - [@video@What is Reverse ETL?](https://www.youtube.com/watch?v=DRAGfc5or2Y) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/scala@WHJXJ5ukJd-tK_3LFLJBg.md b/src/data/roadmaps/data-engineer/content/scala@WHJXJ5ukJd-tK_3LFLJBg.md index c5df29fb7..4bae594ba 100644 --- a/src/data/roadmaps/data-engineer/content/scala@WHJXJ5ukJd-tK_3LFLJBg.md +++ b/src/data/roadmaps/data-engineer/content/scala@WHJXJ5ukJd-tK_3LFLJBg.md @@ -4,6 +4,7 @@ Scala is a programming language that combines the strengths of object-oriented a Visit the following resources to learn more: +- [@roadmap@Visit the Dedicated Scala Roadmap](https://roadmap.sh/scala) - [@official@The Scala Programming Language](https://www.scala-lang.org/) - [@article@Scala for Beginners: An Introduction](https://daily.dev/blog/scala-for-beginners-an-introduction) - [@video@Scala Tutorial](https://www.youtube.com/playlist?list=PLS1QulWo1RIagob5D6kMIAvu7DQC5VTh3) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/sentry@i54fx-NV6nWzQVCdi0aKL.md b/src/data/roadmaps/data-engineer/content/sentry@i54fx-NV6nWzQVCdi0aKL.md index 8ee0ebf0d..6d992c19d 100644 --- a/src/data/roadmaps/data-engineer/content/sentry@i54fx-NV6nWzQVCdi0aKL.md +++ b/src/data/roadmaps/data-engineer/content/sentry@i54fx-NV6nWzQVCdi0aKL.md @@ -4,5 +4,4 @@ Sentry tracks your software performance, measuring metrics like throughput and l Visit the following resources to learn more: -- [@official@Sentry](https://sentry.io) - [@official@Sentry Documentation](https://docs.sentry.io/) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/skills-and-responsibilities@3BxbkrBp8veZj38zdwN8s.md b/src/data/roadmaps/data-engineer/content/skills-and-responsibilities@3BxbkrBp8veZj38zdwN8s.md index e86277afc..3fc33cc50 100644 --- a/src/data/roadmaps/data-engineer/content/skills-and-responsibilities@3BxbkrBp8veZj38zdwN8s.md +++ b/src/data/roadmaps/data-engineer/content/skills-and-responsibilities@3BxbkrBp8veZj38zdwN8s.md @@ -1,29 +1,8 @@ # Skills and Responsibilities - -Here’s a list of essential data engineering skills: - -1. SQL & Database Management: Ability to query, manipulate, and design relational databases efficiently using SQL. This is the bread-and-butter for extracting, transforming, and analyzing data. - -2. Data Modeling: Designing schemas and structures (star, snowflake, normalized forms) to optimize storage, performance, and usability of data. - -3. ETL/ELT Development: Building Extract-Transform-Load (or Load-Transform) pipelines to move and reshape data between systems while ensuring quality and consistency. - -4. Big Data Frameworks: Proficiency with tools like Apache Spark, Hadoop, or Flink to process and analyze massive datasets in distributed environments. - -5. Cloud Platforms: Working knowledge of AWS, Azure, or GCP for storage, compute, and orchestration (e.g., S3, BigQuery, Dataflow, Redshift). - -6. Data Warehousing: Understanding concepts and tools (Snowflake, BigQuery, Redshift) for centralizing, optimizing, and querying large volumes of business data. - -7. Workflow Orchestration: Using tools like Apache Airflow, Prefect, or Dagster to automate and schedule complex data pipelines reliably. - -8. Scripting & Programming: Strong skills in Python or Scala for building data processing scripts, automation tasks, and integration with APIs. - -9. Data Governance & Security: Applying practices for data quality, lineage tracking, access control, compliance (GDPR, HIPAA), and encryption. - -10. Monitoring & Performance Optimization: Setting up alerts, logging, and tuning pipelines to ensure they run efficiently, catch errors early, and scale smoothly. + +A data engineer works across a broad set of tools and systems: programming languages, databases, cloud platforms, pipeline orchestration, and distributed computing. Core responsibilities include building and maintaining data pipelines, managing database schemas, optimizing query performance, and ensuring data quality. Collaboration with data scientists, analysts, and software engineers is also a regular part of the role. Visit the following resources to learn more: - [@article@Top Data Engineer Skills and Responsibilities](https://www.simplilearn.com/data-engineer-role-article) -- [@article@5 Essential Data Engineering Skills For 2025](https://www.datacamp.com/blog/essential-data-engineering-skills) - [@video@What skills do you need as a Data Engineer?](https://www.youtube.com/watch?v=sF04UxNAvmg) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/slowly-changing-dimension---scd@5KgPfywItqLFQRnIZldZH.md b/src/data/roadmaps/data-engineer/content/slowly-changing-dimension---scd@5KgPfywItqLFQRnIZldZH.md index acae97064..80fcd420c 100644 --- a/src/data/roadmaps/data-engineer/content/slowly-changing-dimension---scd@5KgPfywItqLFQRnIZldZH.md +++ b/src/data/roadmaps/data-engineer/content/slowly-changing-dimension---scd@5KgPfywItqLFQRnIZldZH.md @@ -4,5 +4,4 @@ Slowly Changing Dimensions (SCDs) are a data warehousing technique used to track Visit the following resources to learn more: -- [@article@WMastering Slowly Changing Dimensions (SCD)](https://www.datacamp.com/tutorial/mastering-slowly-changing-dimensions-scd) - [@article@Implementing Slowly Changing Dimensions (SCDs) in Data Warehouses](https://www.sqlshack.com/implementing-slowly-changing-dimensions-scds-in-data-warehouses/) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/smoke-testing@woa5K4Dt9L6aBzlJMNS31.md b/src/data/roadmaps/data-engineer/content/smoke-testing@woa5K4Dt9L6aBzlJMNS31.md index 5caadea76..7db305699 100644 --- a/src/data/roadmaps/data-engineer/content/smoke-testing@woa5K4Dt9L6aBzlJMNS31.md +++ b/src/data/roadmaps/data-engineer/content/smoke-testing@woa5K4Dt9L6aBzlJMNS31.md @@ -4,5 +4,4 @@ Smoke Testing is a software testing process that determines whether the deployed Visit the following resources to learn more: -- [@article@Smoke Testing | Software Testing](https://www.guru99.com/smoke-testing.html) -- [@feed@Explore top posts about Testing](https://app.daily.dev/tags/testing?ref=roadmapsh) \ No newline at end of file +- [@article@Smoke Testing | Software Testing](https://www.guru99.com/smoke-testing.html) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/snowflake@Pf0_CBGkmSEfWDQ2_iFXr.md b/src/data/roadmaps/data-engineer/content/snowflake@Pf0_CBGkmSEfWDQ2_iFXr.md index a5824924c..3ab392d71 100644 --- a/src/data/roadmaps/data-engineer/content/snowflake@Pf0_CBGkmSEfWDQ2_iFXr.md +++ b/src/data/roadmaps/data-engineer/content/snowflake@Pf0_CBGkmSEfWDQ2_iFXr.md @@ -6,5 +6,4 @@ Visit the following resources to learn more: - [@official@Snowflake Docs](https://docs.snowflake.com/) - [@official@Snowflake in 20 minutes](https://docs.snowflake.com/en/user-guide/tutorials/snowflake-in-20minutes) -- [@article@Snowflake Tutorial For Beginners: From Architecture to Running Databases](https://www.datacamp.com/tutorial/introduction-to-snowflake-for-beginners) - [@video@Learn Snowflake in 2 Hours](https://www.youtube.com/watch?v=mP3QbYURT9k) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/snowflake@W3l1_66fsIqR3MqgBJUmU.md b/src/data/roadmaps/data-engineer/content/snowflake@W3l1_66fsIqR3MqgBJUmU.md index a5824924c..3ab392d71 100644 --- a/src/data/roadmaps/data-engineer/content/snowflake@W3l1_66fsIqR3MqgBJUmU.md +++ b/src/data/roadmaps/data-engineer/content/snowflake@W3l1_66fsIqR3MqgBJUmU.md @@ -6,5 +6,4 @@ Visit the following resources to learn more: - [@official@Snowflake Docs](https://docs.snowflake.com/) - [@official@Snowflake in 20 minutes](https://docs.snowflake.com/en/user-guide/tutorials/snowflake-in-20minutes) -- [@article@Snowflake Tutorial For Beginners: From Architecture to Running Databases](https://www.datacamp.com/tutorial/introduction-to-snowflake-for-beginners) - [@video@Learn Snowflake in 2 Hours](https://www.youtube.com/watch?v=mP3QbYURT9k) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/star-vs-snowflake-schema@OfH_UXnxvGQgwlNQwOEfS.md b/src/data/roadmaps/data-engineer/content/star-vs-snowflake-schema@OfH_UXnxvGQgwlNQwOEfS.md index 355a4d847..a5b80cfad 100644 --- a/src/data/roadmaps/data-engineer/content/star-vs-snowflake-schema@OfH_UXnxvGQgwlNQwOEfS.md +++ b/src/data/roadmaps/data-engineer/content/star-vs-snowflake-schema@OfH_UXnxvGQgwlNQwOEfS.md @@ -1,11 +1,3 @@ # Star vs Snowflake Schema - -A star schema is a way to organize data in a database, namely in data warehouses, to make it easier and faster to analyze. At the center, there's a main table called the **fact table**, which holds measurable data like sales or revenue. Around it are **dimension tables**, which add details like product names, customer info, or dates. This layout forms a star-like shape. - -A snowflake schema is another way of organizing data. In this schema, dimension tables are split into smaller sub-dimensions to keep data more organized and detailed, just like snowflakes in a large lake. - -The star schema is simple and fast -ideal when you need to extract data for analysis quickly. On the other hand, the snowflake schema is more detailed. It prioritizes storage efficiency and managing complex data relationships. - -Visit the following resources to learn more: - -- [@official@Star Schema vs Snowflake Schema: Differences & Use Cases](https://www.datacamp.com/blog/star-schema-vs-snowflake-schema) \ No newline at end of file + +Star and snowflake schemas are two approaches to organizing data in a data warehouse. A star schema has a central fact table connected directly to dimension tables, making queries simple and fast. A snowflake schema normalizes dimension tables into multiple related tables, reducing redundancy but requiring more joins. Star schemas are more common in analytical systems due to their query performance. \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/streamlit@FfU6Vwf0PXva91FoqxFgp.md b/src/data/roadmaps/data-engineer/content/streamlit@FfU6Vwf0PXva91FoqxFgp.md index 0933c99a4..4152f5736 100644 --- a/src/data/roadmaps/data-engineer/content/streamlit@FfU6Vwf0PXva91FoqxFgp.md +++ b/src/data/roadmaps/data-engineer/content/streamlit@FfU6Vwf0PXva91FoqxFgp.md @@ -5,5 +5,4 @@ Streamlit is a free and open-source framework to rapidly build and share machine Visit the following resources to learn more: - [@official@Streamlit Docs](https://docs.streamlit.io/) -- [@official@Streamlit Python: Tutorial](https://www.datacamp.com/tutorial/streamlit) - [@video@EStreamlit Explained: Python Tutorial for Data Scientists](https://www.youtube.com/watch?v=c8QXUrvSSyg) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/testing@DZoxLu-j1vq5leoXLRZqt.md b/src/data/roadmaps/data-engineer/content/testing@DZoxLu-j1vq5leoXLRZqt.md index 53dc7ede3..c1a459663 100644 --- a/src/data/roadmaps/data-engineer/content/testing@DZoxLu-j1vq5leoXLRZqt.md +++ b/src/data/roadmaps/data-engineer/content/testing@DZoxLu-j1vq5leoXLRZqt.md @@ -1,9 +1,8 @@ # Testing - -Testing is a systematic process used to evaluate the functionality, performance, and quality of software or systems to ensure they meet specified requirements and standards. It involves various methodologies and levels, including unit testing (testing individual components), integration testing (verifying interactions between components), system testing (assessing the entire system's behavior), and acceptance testing (confirming it meets user needs). Testing can be manual or automated and aims to identify defects, validate that features work as intended, and ensure the system performs reliably under different conditions. Effective testing is critical for delivering high-quality software and mitigating risks before deployment. + +Testing in data engineering involves verifying that pipelines, transformations, and data outputs behave correctly. This includes unit tests for individual functions, integration tests for pipeline components, and data quality tests that validate the output data itself. A well-tested pipeline catches regressions early and builds confidence in the reliability of data delivered to consumers. Visit the following resources to learn more: - [@article@What is Software Testing?](https://www.guru99.com/software-testing-introduction-importance.html) -- [@article@Testing Pyramid](https://www.browserstack.com/guide/testing-pyramid-for-test-automation) -- [@feed@Explore top posts about Testing](https://app.daily.dev/tags/testing?ref=roadmapsh) \ No newline at end of file +- [@article@Testing Pyramid](https://www.browserstack.com/guide/testing-pyramid-for-test-automation) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/tokenization@ZAKo9Svb8TQ6KkmOnfB5x.md b/src/data/roadmaps/data-engineer/content/tokenization@ZAKo9Svb8TQ6KkmOnfB5x.md index 1f2675fb2..aa7c6254b 100644 --- a/src/data/roadmaps/data-engineer/content/tokenization@ZAKo9Svb8TQ6KkmOnfB5x.md +++ b/src/data/roadmaps/data-engineer/content/tokenization@ZAKo9Svb8TQ6KkmOnfB5x.md @@ -1,8 +1,7 @@ # Tokenization - -Tokenization is the step where raw text is broken into small pieces called tokens, and each token is given a unique number. A token can be a whole word, part of a word, a punctuation mark, or even a space. The list of all possible tokens is the model’s vocabulary. Once text is turned into these numbered tokens, the model can look up an embedding for each number and start its math. By working with tokens instead of full sentences, the model keeps the input size steady and can handle new or rare words by slicing them into familiar sub-pieces. After the model finishes its work, the numbered tokens are turned back into text through the same vocabulary map, letting the user read the result. + +Tokenization replaces sensitive data values with non-sensitive placeholders called tokens. The original value is stored securely in a token vault, and the token can be used in systems that do not need the actual data. Tokenization is used to protect payment card numbers, personal identifiers, and other sensitive field Visit the following resources to learn more: -- [@article@Explaining Tokens — the Language and Currency of AI](https://blogs.nvidia.com/blog/ai-tokens-explained/) -- [@article@What is Tokenization? Types, Use Cases, Implementation](https://www.datacamp.com/blog/what-is-tokenization) \ No newline at end of file +- [@article@Explaining Tokens — the Language and Currency of AI](https://blogs.nvidia.com/blog/ai-tokens-explained/) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/unit-testing@8dXD4ddR_USEbAJhUMcB6.md b/src/data/roadmaps/data-engineer/content/unit-testing@8dXD4ddR_USEbAJhUMcB6.md index 5b395b83a..2c32eae1e 100644 --- a/src/data/roadmaps/data-engineer/content/unit-testing@8dXD4ddR_USEbAJhUMcB6.md +++ b/src/data/roadmaps/data-engineer/content/unit-testing@8dXD4ddR_USEbAJhUMcB6.md @@ -5,5 +5,4 @@ Unit testing is where individual **units** (modules, functions/methods, routines Visit the following resources to learn more: - [@article@Unit Testing Tutorial](https://www.guru99.com/unit-testing-guide.html) -- [@video@What is Unit Testing?](https://youtu.be/3kzHmaeozDI) -- [@feed@Explore top posts about Testing](https://app.daily.dev/tags/testing?ref=roadmapsh) \ No newline at end of file +- [@video@What is Unit Testing?](https://youtu.be/3kzHmaeozDI) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/what-and-why-use-them@1qju7UlcMo2Ebp4a3BGxH.md b/src/data/roadmaps/data-engineer/content/what-and-why-use-them@1qju7UlcMo2Ebp4a3BGxH.md index e96e85f85..de695fda1 100644 --- a/src/data/roadmaps/data-engineer/content/what-and-why-use-them@1qju7UlcMo2Ebp4a3BGxH.md +++ b/src/data/roadmaps/data-engineer/content/what-and-why-use-them@1qju7UlcMo2Ebp4a3BGxH.md @@ -1,3 +1,3 @@ # What and why use them? - -In data engineering, messaging systems act as central brokers for data communication, allowing different applications and services to send and receive data in a decoupled, scalable, and fault-tolerant way. They are crucial for handling high-volume, real-time data streams, building resilient data pipelines, and enabling event-driven architectures by acting as buffers and communication channels between data producers and consumers. Key benefits include decoupling systems for agility, ensuring data reliability through queuing and retries, and horizontal scalability to manage growing data loads, while common examples include Apache Kafka and message queues like RabbitMQ and AWS SQS. \ No newline at end of file + +Messaging systems solve the problem of tight coupling between systems. Instead of one service directly calling another, it sends a message to a broker, and the consumer reads it when ready. This improves reliability, scalability, and flexibility, especially when producers and consumers operate at different speeds or scales. \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/what-is-cluster-computing@Ad10evrGQuYRl5GaMhQwu.md b/src/data/roadmaps/data-engineer/content/what-is-cluster-computing@Ad10evrGQuYRl5GaMhQwu.md index 4e5394b1b..65c916715 100644 --- a/src/data/roadmaps/data-engineer/content/what-is-cluster-computing@Ad10evrGQuYRl5GaMhQwu.md +++ b/src/data/roadmaps/data-engineer/content/what-is-cluster-computing@Ad10evrGQuYRl5GaMhQwu.md @@ -1,12 +1,9 @@ # What is Cluster Computing - -Cluster computing is a type of distributing computing where multiple computers are connected so they work together as a single system. By working together, a cluster of machines can address complex tasks with higher computational power and efficiency. - -The term “cluster” refers to the network of linked computer systems programmed to perform the same task. Computing clusters typically consist of servers, workstations and personal computers (PCs) that communicate over a local area network (LAN) or a wide area network (WAN). Each computer or “node,” in a computer network has an operating system (OS) and a central processing unit (CPU) core that handles the tasks required for the software to run properly. + +Cluster computing is a model where multiple computers are networked together to act as a unified processing system. Tasks are split across nodes in the cluster and executed in parallel. This approach is used in big data processing, scientific computing, and any workload that exceeds single-machine capacity. Visit the following resources to learn more: - [@article@What is cluster computing? - IBM](https://www.ibm.com/think/topics/cluster-computing) -- [@article@What is cluster computing? - AWS](https://aws.amazon.com/what-is/cluster-computing/) - [@article@Computer cluster - Wikipedia](http://en.wikipedia.org/wiki/Computer_cluster) - [@video@WUnderstand the Basic Cluster Concepts](https://www.youtube.com/watch?v=8BBDxzJL6fY) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/what-is-data-engineering@WB2PRVI9C6RIbJ6l9zdbd.md b/src/data/roadmaps/data-engineer/content/what-is-data-engineering@WB2PRVI9C6RIbJ6l9zdbd.md index 48c405902..559811f6f 100644 --- a/src/data/roadmaps/data-engineer/content/what-is-data-engineering@WB2PRVI9C6RIbJ6l9zdbd.md +++ b/src/data/roadmaps/data-engineer/content/what-is-data-engineering@WB2PRVI9C6RIbJ6l9zdbd.md @@ -5,5 +5,4 @@ Data engineering is the practice of designing and building systems for the aggre Visit the following resources to learn more: - [@article@What is data engineering?](https://www.ibm.com/think/topics/data-engineering) -- [@article@How to Become a Data Engineer in 2025: 5 Steps for Career Success](https://www.datacamp.com/blog/how-to-become-a-data-engineer) - [@video@How Data Engineering Works?](https://www.youtube.com/watch?v=qWru-b6m030) \ No newline at end of file diff --git a/src/data/roadmaps/data-engineer/content/what-is-data-warehouse@dc3lJI27hJ3zZ45UCVqM1.md b/src/data/roadmaps/data-engineer/content/what-is-data-warehouse@dc3lJI27hJ3zZ45UCVqM1.md index 6aef6f35c..79d655f6d 100644 --- a/src/data/roadmaps/data-engineer/content/what-is-data-warehouse@dc3lJI27hJ3zZ45UCVqM1.md +++ b/src/data/roadmaps/data-engineer/content/what-is-data-warehouse@dc3lJI27hJ3zZ45UCVqM1.md @@ -1,6 +1,6 @@ -# Data Warehouse - -**Data Warehouses** are data storage systems which are designed for analyzing, reporting and integrating with transactional systems. The data in a warehouse is clean, consistent, and often transformed to meet wide-range of business requirements. Hence, data warehouses provide structured data but require more processing and management compared to data lakes. +# What is Data Warehouse? + +A data warehouse is a centralized repository for storing large volumes of structured, historical data from multiple sources. It is optimized for analytical queries rather than transactional operations. Data warehouses power business intelligence, reporting, and data analysis, providing a single source of truth across an organization. Visit the following resources to learn more: