From e2f80a51beab594ae2b95e019797354b35558c9b Mon Sep 17 00:00:00 2001 From: "github-actions[bot]" <41898282+github-actions[bot]@users.noreply.github.com> Date: Tue, 16 Jun 2026 13:12:17 +0200 Subject: [PATCH] chore: sync content to repo (#10082) Co-authored-by: nilbuild <4921183+nilbuild@users.noreply.github.com> --- .../content/airflow@LERo7-ggtfZ-KZyNdBwt6.md | 10 ++++++++-- .../content/altair@9vgD1Sd5B3iRMPLrfMG0C.md | 2 +- .../apis-with-requests@10c4tiBOR68J3qa55LcJS.md | 10 ++++++++-- .../content/args--kwargs@7ukR74cN6lHbvhs4mFn6R.md | 9 +++++++-- .../content/arithmetic@AAxvIAas2Sdkz9-2uNHXq.md | 11 +++++++++-- .../content/array-operations@qOZ-T7zaPjvIT222rSV4o.md | 9 +++++++-- .../content/arrays--ndarray@qiUPTXFgfukxlPhnv76jo.md | 10 ++++++++-- .../content/beautifulsoup@w4K_ffCUbzhGxWyoKsspu.md | 10 ++++++++-- .../content/big-data-tools@GkxSrpOJx-SWvnDkTPcJt.md | 8 ++++++-- .../content/booleans@Ze02ApxTvyQqtB3VYv3d4.md | 9 +++++++-- .../content/boxplot@ij8PMz519gC9WIxlWGD7B.md | 10 ++++++++-- .../built-in-functions@KtV5Jl-hOcUS5L-15BUDD.md | 9 +++++++-- .../content/casting-types@XI8bvhQ3ugezTu60USUG2.md | 9 +++++++-- .../categorical-plots@Dj8Ar4-NQ5mQQhNHmQJVy.md | 8 ++++++-- .../content/comparison@Md0yXC3JTH56Wky5jA8Fk.md | 9 +++++++-- .../content/conda@j1jq5_64lBuYQQjJWKIuI.md | 10 ++++++++-- .../content/conditionals@NP1kjSk0ujU0Gx-ajNHlR.md | 10 ++++++++-- .../correlation--covariance@YCMSjDjUCnkz2VPXfLxz5.md | 9 +++++++-- .../correlation-matrix@PHfLEgeEMChsSOSIgMoWH.md | 9 +++++++-- .../content/cross-tabulation@7XHYCVpUpvwAtaUBjTa4O.md | 10 ++++++++-- .../content/csv@nDUOJPCEzkCxX8B63HBP3.md | 10 ++++++++-- .../customizing-plots@JTYST4p4O9RQbCzxsoQhA.md | 10 ++++++++-- .../customizing-plots@j4UHICw-niFv1RRlC-vk1.md | 9 +++++++-- .../content/dash@-dsafcJlA16H_ji0zQvKy.md | 9 +++++++-- .../content/dashboards@uyhlD_ZBCCd3IvrWZ5mlh.md | 9 +++++++-- .../content/dask@ApowKoA3C1B8oUr4-sltZ.md | 9 +++++++-- .../content/data-cleaning@K53X-NB1S9XXSF6j8q1qu.md | 10 ++++++++-- .../content/data-pipelines@f3YL32h2gbU_G8kkNC3Vc.md | 9 +++++++-- .../defining-functions@0XtKHqPS6AfhQvhH5cu7Q.md | 9 +++++++-- .../content/dictionaries@lReYQM-p8rrYcHkaqzbMr.md | 10 ++++++++-- .../distribution-plots@Xq7Ho2kWg5AtZK0G6z02U.md | 8 ++++++-- .../dropping-vs-imputing@xIS0cJ8ui0siBoZJr3-B_.md | 10 ++++++++-- .../content/duckdb@a6HiqhsJJijqObw9pNOqb.md | 10 ++++++++-- .../encoding-categories@1bHv3_acNts_BeOn3l6DF.md | 9 +++++++-- .../environment-setup@8Z20apU8ugXwdH8rRZAcr.md | 8 ++++++-- .../content/excel@OwEBYI4u7nb9-UGOR9SQf.md | 9 +++++++-- ...exploratory-data-analysis@eObhSoXcbqLdwegzaWmi7.md | 10 ++++++++-- .../filtering--querying@wxO7f8wh6JTXHYDiPcbmy.md | 11 +++++++++-- .../content/floats@aLIXcqB258WNrdASuN_7G.md | 9 +++++++-- .../forward--backward-fill@9mFT0z5yT446_-nGdG492.md | 10 ++++++++-- .../functions--methods@kyrS5ez7_85AnNt8Bojz_.md | 10 ++++++++-- .../content/google-colab@c8uLNJMr6UZfvjFwFhxvd.md | 10 ++++++++-- .../groupby--aggregation@KnSv5dPToipoz2RF1dDnO.md | 9 +++++++-- .../content/heatmaps@Bk3z1CHnrnV6ii2Y0BF8G.md | 10 ++++++++-- .../content/histogram@oRaag-FOl50NJM0y4joBm.md | 10 ++++++++-- .../indexing--slicing@PlFcaQsvFSK0Y-LKGu0eH.md | 11 +++++++++-- .../indexing--slicing@mAOljEWt_0Hb9jLBKuc-H.md | 10 ++++++++-- .../content/integers@_-bUq5jtewDmsmnzYw1-q.md | 9 +++++++-- ...interactive-visualization@ES8zVixZ5wZ0tW6beOWxJ.md | 8 ++++++-- .../content/introduction@GISOFMKvnBys0O0IMpz2J.md | 11 +++++++++-- .../content/iqr@ucEFYVtXIvuYKjudK3pwB.md | 10 ++++++++-- .../content/isnull-isna@bcODvfBqCUwLVXa8HKOFB.md | 11 +++++++++-- .../content/json@VEetEvKGNzAO9QdZY2g86.md | 9 +++++++-- .../content/jupyterlab@2m8uQRGaVZUd4mfm7bRou.md | 9 +++++++-- .../content/lambda-functions@AzsZhpK-wuWtafVFJMhdY.md | 9 +++++++-- .../linear-algebra-basics@xHDrwKDE_mrSihCsPQMBF.md | 9 +++++++-- .../list-comprehensions@la9oiXURvw6AH-Vuh-WqT.md | 9 +++++++-- .../content/lists@F4dp3Ip8G-EI1hprtv3do.md | 9 +++++++-- .../content/logical@_0K6o-R0F2a9bpEDfRcTi.md | 10 ++++++++-- .../content/loops@Dvy7BnNzK55qbh_SgOk8m.md | 11 +++++++++-- .../content/matplotlib@Md9Yq7bwDNo6hO_A9aa9N.md | 10 ++++++++-- .../content/mean-median-mode@DdA2gJKpgOHj9mqZL-c4F.md | 9 +++++++-- .../content/merging--joining@Rq9TFEEJnjEgVa1DiJxFz.md | 9 +++++++-- .../content/numpy@IO-disDSfkPau4Pd_TNwc.md | 10 ++++++++-- .../oop-for-data-analysis@hTAQSe0rHv7fuAMfxJwGp.md | 8 +++++++- .../content/operators@so95CO6Qw3I0S98ISENS-.md | 10 ++++++++-- .../pandas-string-methods@yRHodYYqdICLZyG3J9qjc.md | 10 ++++++++-- .../content/pandas@bIblb5cVNSZLqSuskmKnP.md | 10 ++++++++-- .../content/pandas@cD139cRk1xj2ip-o_078l.md | 9 +++++++-- .../content/parquet@sBKhhjfFTpbj7TQwVnX7F.md | 9 +++++++-- .../content/parsing-dates@L7btVRwvkEVX0K8LxnbqJ.md | 10 ++++++++-- .../content/pip@6dJaQRj2zMz00sfTiTdww.md | 10 ++++++++-- .../content/plot-categories@5HEYmbVJ6Z7lvgqalVjxW.md | 10 ++++++++-- .../content/plotly@J_u8yvrHEWKKmVsVgPN7L.md | 10 ++++++++-- .../content/polars@18eitPvBQKdeM3K33d59g.md | 9 +++++++-- .../power-bi--tableau@EHVXO79NGpLC1i8x4Y7g1.md | 9 +++++++-- .../printing-variables@icVUQgedEGcaPpJSV8JdI.md | 11 +++++++++-- .../content/pyspark@XTXm8aIDbRkTXoAck_m3I.md | 10 ++++++++-- .../content/random-module@XizmE-QroU4GJbjvqNYxe.md | 10 ++++++++-- .../content/re@oXhGEIie-MzfUSM2Ng4Zm.md | 11 +++++++++-- .../content/reading-data@w2j11Jyr_qLP1Y9ST-06O.md | 10 ++++++++-- .../reading-local-files@34DL5RWQtHkvgVhI9a_ZQ.md | 9 +++++++-- .../content/reading-web-data@w8iQWthKu1w75z36B8Ir3.md | 8 ++++++-- .../content/regression-plots@PcGJMkmZrl4xc2ADT9dF3.md | 10 ++++++++-- .../content/reshaping@ACr__YweG2xOAf56FK8tE.md | 10 ++++++++-- .../content/saving-figures@mElhFGGT2GnLg2AkJFETN.md | 9 +++++++-- .../content/scatterplot@4ce2yYLRCSb3azsOMpE-U.md | 10 ++++++++-- .../content/scikit-learn@npndQC8OIfP1wwn_iswPy.md | 9 +++++++-- .../content/scipy@YxA0ms9vkufKGOguRNK1B.md | 10 ++++++++-- .../content/scrapy@UImXdBHQKCzp9NdjF9xMi.md | 10 ++++++++-- .../content/seaborn@zx7uCaQiCxvXMhm9FG3XG.md | 11 +++++++++-- .../series-and-dataframe@3N5N6qe8JVr0XGOFdn5-8.md | 11 +++++++++-- .../content/sets@GvV_thqIC0yXyJbJt_Zmu.md | 9 +++++++-- .../content/sql-fundamentals@E_r4BLrSyHkrxpg4HDkUm.md | 10 ++++++++-- .../content/sqlalchemy@k63rFXaVDUX2ZOVI_O9WG.md | 10 ++++++++-- .../content/sqlite3@-RepeaHp66GGuM66k9JHd.md | 9 +++++++-- .../content/statistics--ml@KLXFKVb-EC5WOigI63y09.md | 10 ++++++++-- .../content/streamlit@Z-inS5OKdXi6yveKgxMeC.md | 10 ++++++++-- .../content/strings@F-0Rxwu_RcNF_NOoFHoo7.md | 9 +++++++-- .../strip-replace-split@S1EzGF-CLu7tWDb1V6RQc.md | 9 +++++++-- .../subplots-and-figures@GRg7HyeVuSwkKD5bCOFiY.md | 10 ++++++++-- .../content/tuples@olniprgaL7l-9YhwbFfmN.md | 9 +++++++-- .../content/type-casting@NjCor7ePiZapd4f6bMZlV.md | 9 +++++++-- .../variance--std-deviation@nZl1ngETFxKStt0tFvOpR.md | 9 +++++++-- .../content/virtualenv--venv@IL0fFEs4CgK_eBft-SYAl.md | 10 ++++++++-- .../visual-inspection@UIFNQ9g5fMK-KoZCrqv27.md | 9 +++++++-- .../content/vs-code@ObA_xZDY7PxU54NGBwyVI.md | 9 +++++++-- .../working-with-strings@Sg5w8zO2Ji-uDJKEoWey9.md | 10 ++++++++-- .../content/z-score@2HIeG6ywOA9BkE9W3Gu9v.md | 9 +++++++-- 109 files changed, 818 insertions(+), 216 deletions(-) diff --git a/src/data/roadmaps/python-data-analysis/content/airflow@LERo7-ggtfZ-KZyNdBwt6.md b/src/data/roadmaps/python-data-analysis/content/airflow@LERo7-ggtfZ-KZyNdBwt6.md index cb72dbd70..3df024631 100644 --- a/src/data/roadmaps/python-data-analysis/content/airflow@LERo7-ggtfZ-KZyNdBwt6.md +++ b/src/data/roadmaps/python-data-analysis/content/airflow@LERo7-ggtfZ-KZyNdBwt6.md @@ -1,3 +1,9 @@ # Airflow - -Apache Airflow is an open-source platform for authoring, scheduling, and monitoring data pipelines. Pipelines are defined as DAGs (Directed Acyclic Graphs) in Python, where each node is a task and edges define dependencies. Airflow provides a web UI for monitoring runs, retrying failures, and tracking execution history. It is the standard orchestration tool for production Python data pipelines. \ No newline at end of file + +Apache Airflow is an open-source platform for authoring, scheduling, and monitoring data pipelines. Pipelines are defined as DAGs (Directed Acyclic Graphs) in Python, where each node is a task and edges define dependencies. Airflow provides a web UI for monitoring runs, retrying failures, and tracking execution history. It is the standard orchestration tool for production Python data pipelines. + +Visit the following resources to learn more: + +- [@official@Airflow 101: Building Your First Workflow](https://airflow.apache.org/docs/apache-airflow/stable/tutorial/fundamentals.html) +- [@article@Introduction to Apache Airflow](https://www.dataquest.io/blog/introduction-to-apache-airflow/) +- [@video@Airflow Tutorial For Beginners (2026)](https://www.youtube.com/watch?v=IiczxlbQb8s) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/altair@9vgD1Sd5B3iRMPLrfMG0C.md b/src/data/roadmaps/python-data-analysis/content/altair@9vgD1Sd5B3iRMPLrfMG0C.md index fec6b01fd..19bc4498a 100644 --- a/src/data/roadmaps/python-data-analysis/content/altair@9vgD1Sd5B3iRMPLrfMG0C.md +++ b/src/data/roadmaps/python-data-analysis/content/altair@9vgD1Sd5B3iRMPLrfMG0C.md @@ -1,3 +1,3 @@ # Altair - + Altair is a declarative statistical visualization library for Python based on the Vega-Lite grammar. Charts are built by binding data columns to visual channels (x, y, color, size) and specifying the mark type. Altair produces interactive charts by default and generates JSON specifications that render in notebooks and web browsers. \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/apis-with-requests@10c4tiBOR68J3qa55LcJS.md b/src/data/roadmaps/python-data-analysis/content/apis-with-requests@10c4tiBOR68J3qa55LcJS.md index 0f61a9221..70b70a549 100644 --- a/src/data/roadmaps/python-data-analysis/content/apis-with-requests@10c4tiBOR68J3qa55LcJS.md +++ b/src/data/roadmaps/python-data-analysis/content/apis-with-requests@10c4tiBOR68J3qa55LcJS.md @@ -1,3 +1,9 @@ # APIs with requests - -The `requests` library is the standard Python tool for making HTTP requests. It is used to call REST APIs that return JSON or XML data. A typical workflow involves calling `requests.get(url, params=params)`, checking the response status, and parsing the JSON body with `.json()` before loading it into a DataFrame. \ No newline at end of file + +The `requests` library is the standard Python tool for making HTTP requests. It is used to call REST APIs that return JSON or XML data. A typical workflow involves calling `requests.get(url, params=params)`, checking the response status, and parsing the JSON body with `.json()` before loading it into a DataFrame. + +Visit the following resources to learn more: + +- [@article@Python Requests Module](https://www.w3schools.com/python/module_requests.asp) +- [@article@Python's Requests Library (Guide)](https://realpython.com/python-requests/) +- [@video@Master Python Requests In 15 Minutes. Call Any API](https://www.youtube.com/watch?v=Xnbef8F_Yfc) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/args--kwargs@7ukR74cN6lHbvhs4mFn6R.md b/src/data/roadmaps/python-data-analysis/content/args--kwargs@7ukR74cN6lHbvhs4mFn6R.md index 5f0a07a5c..dbba5a08b 100644 --- a/src/data/roadmaps/python-data-analysis/content/args--kwargs@7ukR74cN6lHbvhs4mFn6R.md +++ b/src/data/roadmaps/python-data-analysis/content/args--kwargs@7ukR74cN6lHbvhs4mFn6R.md @@ -1,3 +1,8 @@ # args & kwargs - -`*args` allows a function to accept any number of positional arguments as a tuple. `**kwargs` allows any number of keyword arguments as a dictionary. They make functions flexible when the number or names of arguments are not known in advance, and are widely used in Python libraries for passing options through layers of function calls. \ No newline at end of file + +`*args` allows a function to accept any number of positional arguments as a tuple. `**kwargs` allows any number of keyword arguments as a dictionary. They make functions flexible when the number or names of arguments are not known in advance, and are widely used in Python libraries for passing options through layers of function calls. + +Visit the following resources to learn more: + +- [@article@Python args & kwargs](https://www.w3schools.com/python/python_args_kwargs.asp) +- [@video@Python *ARGS & **KWARGS are awesome!](https://www.youtube.com/watch?v=Vh__2V2tXUM) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/arithmetic@AAxvIAas2Sdkz9-2uNHXq.md b/src/data/roadmaps/python-data-analysis/content/arithmetic@AAxvIAas2Sdkz9-2uNHXq.md index 4bfe8e3b6..d0fadb926 100644 --- a/src/data/roadmaps/python-data-analysis/content/arithmetic@AAxvIAas2Sdkz9-2uNHXq.md +++ b/src/data/roadmaps/python-data-analysis/content/arithmetic@AAxvIAas2Sdkz9-2uNHXq.md @@ -1,3 +1,10 @@ # Arithmetic - -Arithmetic operators perform mathematical calculations: `+` (addition), `-` (subtraction), `*` (multiplication), `/` (division), `//` (floor division), `%` (modulo), and `**` (exponentiation). They are used constantly for computing derived columns, normalizing values, and performing aggregations. \ No newline at end of file + +Arithmetic operators perform mathematical calculations: `+` (addition), `-` (subtraction), `*` (multiplication), `/` (division), `//` (floor division), `%` (modulo), and `**` (exponentiation). They are used constantly for computing derived columns, normalizing values, and performing aggregations. + +Visit the following resources to learn more: + +- [@article@Python Arithmetic Operators](https://www.w3schools.com/python/python_operators_arithmetic.asp) +- [@article@Python Exponent: 5 Methods for Exponentiation + Applications](https://roadmap.sh/python/exponent) +- [@article@Python Division: Operators, Floor Division, and Examples](https://roadmap.sh/python/division) +- [@article@Python Modulo Operator (%): Complete Guide with Examples](https://roadmap.sh/python/modulo) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/array-operations@qOZ-T7zaPjvIT222rSV4o.md b/src/data/roadmaps/python-data-analysis/content/array-operations@qOZ-T7zaPjvIT222rSV4o.md index 67d5bd55f..63851e39c 100644 --- a/src/data/roadmaps/python-data-analysis/content/array-operations@qOZ-T7zaPjvIT222rSV4o.md +++ b/src/data/roadmaps/python-data-analysis/content/array-operations@qOZ-T7zaPjvIT222rSV4o.md @@ -1,3 +1,8 @@ # Array Operations - -NumPy supports a wide range of array operations: element-wise arithmetic, aggregation functions (`sum`, `mean`, `std`, `min`, `max`), reshaping, stacking, and splitting. These operations are vectorized, meaning they apply to the entire array at once without explicit loops, making them highly efficient. \ No newline at end of file + +NumPy supports a wide range of array operations: element-wise arithmetic, aggregation functions (`sum`, `mean`, `std`, `min`, `max`), reshaping, stacking, and splitting. These operations are vectorized, meaning they apply to the entire array at once without explicit loops, making them highly efficient. + +Visit the following resources to learn more: + +- [@article@NumPy Arithmetic Array Operations](https://www.programiz.com/python-programming/numpy/basic-array-operations) +- [@video@Ultimate Guide to NumPy Arrays](https://www.youtube.com/watch?v=lLRBYKwP8GQ) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/arrays--ndarray@qiUPTXFgfukxlPhnv76jo.md b/src/data/roadmaps/python-data-analysis/content/arrays--ndarray@qiUPTXFgfukxlPhnv76jo.md index 4b358f864..304db9b79 100644 --- a/src/data/roadmaps/python-data-analysis/content/arrays--ndarray@qiUPTXFgfukxlPhnv76jo.md +++ b/src/data/roadmaps/python-data-analysis/content/arrays--ndarray@qiUPTXFgfukxlPhnv76jo.md @@ -1,3 +1,9 @@ # Arrays & ndarray - -The `ndarray` is NumPy's core data structure: a multi-dimensional, homogeneously typed array stored in contiguous memory. It supports element-wise operations, broadcasting, and vectorized computation far faster than Python lists. Understanding ndarray is fundamental to working efficiently with numerical data in Python. \ No newline at end of file + +The `ndarray` is NumPy's core data structure: a multi-dimensional, homogeneously typed array stored in contiguous memory. It supports element-wise operations, broadcasting, and vectorized computation far faster than Python lists. Understanding ndarray is fundamental to working efficiently with numerical data in Python. + +Visit the following resources to learn more: + +- [@official@NumPy: the absolute basics for beginners](https://numpy.org/doc/stable/user/absolute_beginners.html) +- [@official@numpy.array](https://numpy.org/doc/stable/reference/generated/numpy.array.html) +- [@video@Ultimate Guide to NumPy Arrays](https://www.youtube.com/watch?v=lLRBYKwP8GQ&pp=ygUMbnVtcHkgYXJyYXlz) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/beautifulsoup@w4K_ffCUbzhGxWyoKsspu.md b/src/data/roadmaps/python-data-analysis/content/beautifulsoup@w4K_ffCUbzhGxWyoKsspu.md index 0d780a1ab..e58208240 100644 --- a/src/data/roadmaps/python-data-analysis/content/beautifulsoup@w4K_ffCUbzhGxWyoKsspu.md +++ b/src/data/roadmaps/python-data-analysis/content/beautifulsoup@w4K_ffCUbzhGxWyoKsspu.md @@ -1,3 +1,9 @@ # BeautifulSoup - -BeautifulSoup is a Python library for parsing HTML and XML documents. It provides methods for navigating the document tree, searching for elements by tag, class, or attribute, and extracting text and links. BeautifulSoup is used for web scraping when the target website does not provide an API. \ No newline at end of file + +BeautifulSoup is a Python library for parsing HTML and XML documents. It provides methods for navigating the document tree, searching for elements by tag, class, or attribute, and extracting text and links. BeautifulSoup is used for web scraping when the target website does not provide an API. + +Visit the following resources to learn more: + +- [@official@Beautiful Soup Documentation](https://beautiful-soup-4.readthedocs.io/en/latest/) +- [@article@Beautiful Soup: Build a Web Scraper With Python](https://realpython.com/beautiful-soup-web-scraper-python/) +- [@video@Web Scraping with Python - Beautiful Soup Crash Course](https://www.youtube.com/watch?v=XVv6mJpFOb0) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/big-data-tools@GkxSrpOJx-SWvnDkTPcJt.md b/src/data/roadmaps/python-data-analysis/content/big-data-tools@GkxSrpOJx-SWvnDkTPcJt.md index a7950374d..796cb1196 100644 --- a/src/data/roadmaps/python-data-analysis/content/big-data-tools@GkxSrpOJx-SWvnDkTPcJt.md +++ b/src/data/roadmaps/python-data-analysis/content/big-data-tools@GkxSrpOJx-SWvnDkTPcJt.md @@ -1,3 +1,7 @@ # Big Data Tools - -Big data tools process datasets too large to fit in a single machine's memory using distributed or out-of-core computation. Dask and PySpark are the primary Python tools for scaling beyond what Pandas and NumPy can handle. They provide familiar DataFrame-like APIs while distributing computation across cores or clusters. \ No newline at end of file + +Big data tools process datasets too large to fit in a single machine's memory using distributed or out-of-core computation. Dask and PySpark are the primary Python tools for scaling beyond what Pandas and NumPy can handle. They provide familiar DataFrame-like APIs while distributing computation across cores or clusters. + +Visit the following resources to learn more: + +- [@article@4 Types of Big Data Technologies (+ Management Tools)](https://www.coursera.org/articles/big-data-technologies) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/booleans@Ze02ApxTvyQqtB3VYv3d4.md b/src/data/roadmaps/python-data-analysis/content/booleans@Ze02ApxTvyQqtB3VYv3d4.md index fd75d4b5d..5576b9e3c 100644 --- a/src/data/roadmaps/python-data-analysis/content/booleans@Ze02ApxTvyQqtB3VYv3d4.md +++ b/src/data/roadmaps/python-data-analysis/content/booleans@Ze02ApxTvyQqtB3VYv3d4.md @@ -1,3 +1,8 @@ # Booleans - -Booleans (`bool`) have two values: `True` and `False`. They are the result of comparison and logical operations and are used to control flow and filter data. \ No newline at end of file + +Booleans (`bool`) have two values: `True` and `False`. They are the result of comparison and logical operations and are used to control flow and filter data. + +Visit the following resources to learn more: + +- [@official@Built-in Types](https://docs.python.org/3/library/stdtypes.html) +- [@article@Python Booleans: Use Truth Values in Your Code](https://realpython.com/ref/builtin-types/str/) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/boxplot@ij8PMz519gC9WIxlWGD7B.md b/src/data/roadmaps/python-data-analysis/content/boxplot@ij8PMz519gC9WIxlWGD7B.md index 3faeb959a..d4999118e 100644 --- a/src/data/roadmaps/python-data-analysis/content/boxplot@ij8PMz519gC9WIxlWGD7B.md +++ b/src/data/roadmaps/python-data-analysis/content/boxplot@ij8PMz519gC9WIxlWGD7B.md @@ -1,3 +1,9 @@ # Boxplot - -A box plot displays the five-number summary of a variable: minimum, Q1, median, Q3, and maximum. The box covers the IQR, a line marks the median, and whiskers extend to the data range. Points beyond the whiskers are plotted individually as potential outliers. Box plots are effective for comparing distributions across groups. \ No newline at end of file + +A box plot displays the five-number summary of a variable: minimum, Q1, median, Q3, and maximum. The box covers the IQR, a line marks the median, and whiskers extend to the data range. Points beyond the whiskers are plotted individually as potential outliers. Box plots are effective for comparing distributions across groups. + +Visit the following resources to learn more: + +- [@official@matplotlib.pyplot.boxplot](https://matplotlib.org/stable/api/_as_gen/matplotlib.pyplot.boxplot.html) +- [@article@Create and customize boxplots with Python’s Matplotlib](https://towardsdatascience.com/create-and-customize-boxplots-with-pythons-matplotlib-to-get-lots-of-insights-from-your-data-d561c9883643/) +- [@video@Seaborn boxplot | Box plot explanation, box plot demo, and how to make a box plot in Python seaborn](https://www.youtube.com/watch?v=Vo-bfTqEFQk) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/built-in-functions@KtV5Jl-hOcUS5L-15BUDD.md b/src/data/roadmaps/python-data-analysis/content/built-in-functions@KtV5Jl-hOcUS5L-15BUDD.md index 97868b571..4d512dc96 100644 --- a/src/data/roadmaps/python-data-analysis/content/built-in-functions@KtV5Jl-hOcUS5L-15BUDD.md +++ b/src/data/roadmaps/python-data-analysis/content/built-in-functions@KtV5Jl-hOcUS5L-15BUDD.md @@ -1,3 +1,8 @@ # Built-in Functions - -Python's built-in functions are available without any imports and cover common operations: `len()`, `sum()`, `min()`, `max()`, `sorted()`, `enumerate()`, `zip()`, `map()`, `filter()`, and others. These functions simplify common tasks and are used constantly alongside data analysis libraries. \ No newline at end of file + +Python's built-in functions are available without any imports and cover common operations: `len()`, `sum()`, `min()`, `max()`, `sorted()`, `enumerate()`, `zip()`, `map()`, `filter()`, and others. These functions simplify common tasks and are used constantly alongside data analysis libraries. + +Visit the following resources to learn more: + +- [@official@Built-in Functions](https://docs.python.org/3/library/functions.html) +- [@video@All 71 built-in Python functions](https://www.youtube.com/watch?v=7Qu_KXc7xSI) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/casting-types@XI8bvhQ3ugezTu60USUG2.md b/src/data/roadmaps/python-data-analysis/content/casting-types@XI8bvhQ3ugezTu60USUG2.md index 28a5708b4..78d0ebab9 100644 --- a/src/data/roadmaps/python-data-analysis/content/casting-types@XI8bvhQ3ugezTu60USUG2.md +++ b/src/data/roadmaps/python-data-analysis/content/casting-types@XI8bvhQ3ugezTu60USUG2.md @@ -1,3 +1,8 @@ # Casting Types - -Casting types in Pandas converts a column from one data type to another using `.astype()`. Common conversions include converting string columns to numeric with `pd.to_numeric()`, converting to datetime with `pd.to_datetime()`, and converting integers to categories. Correct data types are required for accurate calculations and efficient memory use. \ No newline at end of file + +Casting types in Pandas converts a column from one data type to another using `.astype()`. Common conversions include converting string columns to numeric with `pd.to_numeric()`, converting to datetime with `pd.to_datetime()`, and converting integers to categories. Correct data types are required for accurate calculations and efficient memory use. + +Visit the following resources to learn more: + +- [@official@pandas.DataFrame.astype](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.astype.html) +- [@article@How To Change Column Type in Pandas DataFrames](https://towardsdatascience.com/how-to-change-column-type-in-pandas-dataframes-d2a5548888f8/) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/categorical-plots@Dj8Ar4-NQ5mQQhNHmQJVy.md b/src/data/roadmaps/python-data-analysis/content/categorical-plots@Dj8Ar4-NQ5mQQhNHmQJVy.md index 44db0d6ec..9c6801f6b 100644 --- a/src/data/roadmaps/python-data-analysis/content/categorical-plots@Dj8Ar4-NQ5mQQhNHmQJVy.md +++ b/src/data/roadmaps/python-data-analysis/content/categorical-plots@Dj8Ar4-NQ5mQQhNHmQJVy.md @@ -1,3 +1,7 @@ # Categorical Plots - -Seaborn's categorical plots visualize the relationship between a numeric variable and one or more categorical variables. `sns.boxplot()`, `sns.violinplot()`, `sns.barplot()`, `sns.stripplot()`, and `sns.countplot()` cover the main patterns. They are used to compare distributions or averages across groups. \ No newline at end of file + +Seaborn's categorical plots visualize the relationship between a numeric variable and one or more categorical variables. `sns.boxplot()`, `sns.violinplot()`, `sns.barplot()`, `sns.stripplot()`, and `sns.countplot()` cover the main patterns. They are used to compare distributions or averages across groups. + +Visit the following resources to learn more: + +- [@official@Visualizing categorical data](https://seaborn.pydata.org/tutorial/categorical.html) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/comparison@Md0yXC3JTH56Wky5jA8Fk.md b/src/data/roadmaps/python-data-analysis/content/comparison@Md0yXC3JTH56Wky5jA8Fk.md index 12f54ea97..f38c740d2 100644 --- a/src/data/roadmaps/python-data-analysis/content/comparison@Md0yXC3JTH56Wky5jA8Fk.md +++ b/src/data/roadmaps/python-data-analysis/content/comparison@Md0yXC3JTH56Wky5jA8Fk.md @@ -1,3 +1,8 @@ # Comparison - -Comparison operators evaluate the relationship between two values and return a boolean. They include `==`, `!=`, `<`, `>`, `<=`, and `>=`. They form the building blocks of filtering conditions applied to DataFrames and arrays. \ No newline at end of file + +Comparison operators evaluate the relationship between two values and return a boolean. They include `==`, `!=`, `<`, `>`, `<=`, and `>=`. They form the building blocks of filtering conditions applied to DataFrames and arrays. + +Visit the following resources to learn more: + +- [@article@Python Comparison Operators](https://www.w3schools.com/python/gloss_python_comparison_operators.asp) +- [@video@Comparison Operators in Python](https://www.youtube.com/watch?v=6ZQtBK-dM9c) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/conda@j1jq5_64lBuYQQjJWKIuI.md b/src/data/roadmaps/python-data-analysis/content/conda@j1jq5_64lBuYQQjJWKIuI.md index b886c0659..b50280b46 100644 --- a/src/data/roadmaps/python-data-analysis/content/conda@j1jq5_64lBuYQQjJWKIuI.md +++ b/src/data/roadmaps/python-data-analysis/content/conda@j1jq5_64lBuYQQjJWKIuI.md @@ -1,3 +1,9 @@ # conda - -conda is an open-source package and environment manager included with the Anaconda and Miniconda distributions. It manages both Python packages and non-Python dependencies, making it well suited for scientific computing. conda environments isolate project dependencies and can be exported to `environment.yml` for reproducibility. \ No newline at end of file + +conda is an open-source package and environment manager included with the Anaconda and Miniconda distributions. It manages both Python packages and non-Python dependencies, making it well suited for scientific computing. conda environments isolate project dependencies and can be exported to `environment.yml` for reproducibility. + +Visit the following resources to learn more: + +- [@official@Anaconda](https://www.anaconda.com/) +- [@video@Anaconda (Conda) for Python - What & Why?](https://www.youtube.com/watch?v=23aQdrS58e0) +- [@video@Master the basics of Conda environments in Python](https://www.youtube.com/watch?v=1VVCd0eSkYc) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/conditionals@NP1kjSk0ujU0Gx-ajNHlR.md b/src/data/roadmaps/python-data-analysis/content/conditionals@NP1kjSk0ujU0Gx-ajNHlR.md index 7c916db05..31bc8751e 100644 --- a/src/data/roadmaps/python-data-analysis/content/conditionals@NP1kjSk0ujU0Gx-ajNHlR.md +++ b/src/data/roadmaps/python-data-analysis/content/conditionals@NP1kjSk0ujU0Gx-ajNHlR.md @@ -1,3 +1,9 @@ # Conditionals - -Conditionals execute different code paths based on whether a condition is true. Python uses `if`, `elif`, and `else` for this. Python 3.10 introduced `match/case`, a structural pattern matching statement that cleanly handles multiple specific value checks as an alternative to long `elif` chains. They appear in custom functions applied to DataFrames, in filtering logic, and in branching pipeline code. \ No newline at end of file + +Conditionals execute different code paths based on whether a condition is true. Python uses `if`, `elif`, and `else` for this. Python 3.10 introduced `match/case`, a structural pattern matching statement that cleanly handles multiple specific value checks as an alternative to long `elif` chains. They appear in custom functions applied to DataFrames, in filtering logic, and in branching pipeline code. + +Visit the following resources to learn more: + +- [@article@Conditional Statements in Python](https://realpython.com/python-conditional-statements/) +- [@article@Python Switch Statement 101: Match-case and alternatives](https://roadmap.sh/python/switch) +- [@video@Control Flow in Python - If Elif Else Statements](https://www.youtube.com/watch?v=Zp5MuPOtsSY) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/correlation--covariance@YCMSjDjUCnkz2VPXfLxz5.md b/src/data/roadmaps/python-data-analysis/content/correlation--covariance@YCMSjDjUCnkz2VPXfLxz5.md index 5a40d2d37..dc944b3a5 100644 --- a/src/data/roadmaps/python-data-analysis/content/correlation--covariance@YCMSjDjUCnkz2VPXfLxz5.md +++ b/src/data/roadmaps/python-data-analysis/content/correlation--covariance@YCMSjDjUCnkz2VPXfLxz5.md @@ -1,3 +1,8 @@ # Correlation & Covariance - -Correlation measures the strength and direction of the linear relationship between two variables, scaled between βˆ’1 and +1. Covariance measures the same relationship but is not normalized, making it harder to interpret across variables with different scales. `df.corr()` and `df.cov()` compute these matrices in Pandas, and heatmaps are used to visualize them. \ No newline at end of file + +Correlation measures the strength and direction of the linear relationship between two variables, scaled between βˆ’1 and +1. Covariance measures the same relationship but is not normalized, making it harder to interpret across variables with different scales. `df.corr()` and `df.cov()` compute these matrices in Pandas, and heatmaps are used to visualize them. + +Visit the following resources to learn more: + +- [@article@Statistics in Python – Understanding Variance, Covariance, and Correlation](https://towardsdatascience.com/statistics-in-python-understanding-variance-covariance-and-correlation-4729b528db01/) +- [@video@Covariance and Correlation in Probability](https://www.youtube.com/watch?v=QKPdk57y7Ck) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/correlation-matrix@PHfLEgeEMChsSOSIgMoWH.md b/src/data/roadmaps/python-data-analysis/content/correlation-matrix@PHfLEgeEMChsSOSIgMoWH.md index 739605b73..e6193f9ce 100644 --- a/src/data/roadmaps/python-data-analysis/content/correlation-matrix@PHfLEgeEMChsSOSIgMoWH.md +++ b/src/data/roadmaps/python-data-analysis/content/correlation-matrix@PHfLEgeEMChsSOSIgMoWH.md @@ -1,3 +1,8 @@ # Correlation Matrix - -A correlation matrix shows the pairwise correlation coefficients between all numeric columns in a dataset. It is computed with `df.corr()` and typically visualized as a heatmap using Seaborn. It is a key EDA tool for identifying which features are strongly related, which helps with feature selection and multicollinearity detection. \ No newline at end of file + +A correlation matrix shows the pairwise correlation coefficients between all numeric columns in a dataset. It is computed with `df.corr()` and typically visualized as a heatmap using Seaborn. It is a key EDA tool for identifying which features are strongly related, which helps with feature selection and multicollinearity detection. + +Visit the following resources to learn more: + +- [@article@Correlation Matrix, Demystified](https://medium.com/data-science/correlation-matrix-demystified-3ae3405c86c1) +- [@article@Data Science - Statistics Correlation Matrix](https://www.w3schools.com/datascience/ds_stat_correlation_matrix.asp) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/cross-tabulation@7XHYCVpUpvwAtaUBjTa4O.md b/src/data/roadmaps/python-data-analysis/content/cross-tabulation@7XHYCVpUpvwAtaUBjTa4O.md index 93fbd0bfc..ae986ef1f 100644 --- a/src/data/roadmaps/python-data-analysis/content/cross-tabulation@7XHYCVpUpvwAtaUBjTa4O.md +++ b/src/data/roadmaps/python-data-analysis/content/cross-tabulation@7XHYCVpUpvwAtaUBjTa4O.md @@ -1,3 +1,9 @@ # Cross-tabulation - -Cross-tabulation (crosstab) counts the frequency of combinations of values across two or more categorical variables. `pd.crosstab()` produces a table of frequencies or proportions. It is used to examine relationships between categorical variables, such as how customer segments differ across product categories. \ No newline at end of file + +Cross-tabulation (crosstab) counts the frequency of combinations of values across two or more categorical variables. `pd.crosstab()` produces a table of frequencies or proportions. It is used to examine relationships between categorical variables, such as how customer segments differ across product categories. + +Visit the following resources to learn more: + +- [@official@pandas.crosstab](https://pandas.pydata.org/docs/reference/api/pandas.crosstab.html) +- [@article@The Power of Crosstab Function in Pandas](https://medium.com/geekculture/the-power-of-crosstab-function-in-pandas-for-data-analysis-and-visualization-6c085c269fcd) +- [@video@Python Pandas Tutorial 13. Crosstab](https://www.youtube.com/watch?v=I_kUj-MfYys) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/csv@nDUOJPCEzkCxX8B63HBP3.md b/src/data/roadmaps/python-data-analysis/content/csv@nDUOJPCEzkCxX8B63HBP3.md index ccfc738cd..412647333 100644 --- a/src/data/roadmaps/python-data-analysis/content/csv@nDUOJPCEzkCxX8B63HBP3.md +++ b/src/data/roadmaps/python-data-analysis/content/csv@nDUOJPCEzkCxX8B63HBP3.md @@ -1,3 +1,9 @@ # CSV - -CSV (Comma-Separated Values) is the most common format for tabular data exchange. `pd.read_csv()` loads a CSV file into a DataFrame and accepts dozens of parameters for handling separators, missing values, date parsing, and data types. CSV files are human-readable but lack type information, so columns often need type correction after loading. \ No newline at end of file + +CSV (Comma-Separated Values) is the most common format for tabular data exchange. `pd.read_csv()` loads a CSV file into a DataFrame and accepts dozens of parameters for handling separators, missing values, date parsing, and data types. CSV files are human-readable but lack type information, so columns often need type correction after loading. + +Visit the following resources to learn more: + +- [@official@pandas.read_csv](https://pandas.pydata.org/docs/reference/api/pandas.read_csv.html) +- [@article@Pandas Read CSV](https://www.w3schools.com/python/pandas/pandas_csv.asp) +- [@video@How to Read a CSV file into a Pandas DataFrame](https://www.youtube.com/watch?v=4YI9z1qUpew) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/customizing-plots@JTYST4p4O9RQbCzxsoQhA.md b/src/data/roadmaps/python-data-analysis/content/customizing-plots@JTYST4p4O9RQbCzxsoQhA.md index b37b9716f..e74f0529b 100644 --- a/src/data/roadmaps/python-data-analysis/content/customizing-plots@JTYST4p4O9RQbCzxsoQhA.md +++ b/src/data/roadmaps/python-data-analysis/content/customizing-plots@JTYST4p4O9RQbCzxsoQhA.md @@ -1,3 +1,9 @@ # Customizing Plots - -Matplotlib allows extensive customization: titles with `set_title()`, axis labels with `set_xlabel()` / `set_ylabel()`, tick formatting, color palettes, line styles, font sizes, legends, and annotations. Customizing plots ensures they communicate clearly and meet the standards required for reports and presentations. \ No newline at end of file + +Matplotlib allows extensive customization: titles with `set_title()`, axis labels with `set_xlabel()` / `set_ylabel()`, tick formatting, color palettes, line styles, font sizes, legends, and annotations. Customizing plots ensures they communicate clearly and meet the standards required for reports and presentations. + +Visit the following resources to learn more: + +- [@course@Advanced Matplotlib: Design & Customize Visualizations](https://www.coursera.org/learn/advanced-matplotlib-design-customize-visualizations) +- [@article@Customizing Plots](https://apxml.com/courses/intermediate-python-programming-ml/chapter-4-data-visualization-matplotlib-seaborn/matplotlib-customization) +- [@video@Matplotlib customization is easy! 🎨](https://www.youtube.com/watch?v=hunq_UOdmoo) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/customizing-plots@j4UHICw-niFv1RRlC-vk1.md b/src/data/roadmaps/python-data-analysis/content/customizing-plots@j4UHICw-niFv1RRlC-vk1.md index 292140908..be0236151 100644 --- a/src/data/roadmaps/python-data-analysis/content/customizing-plots@j4UHICw-niFv1RRlC-vk1.md +++ b/src/data/roadmaps/python-data-analysis/content/customizing-plots@j4UHICw-niFv1RRlC-vk1.md @@ -1,3 +1,8 @@ # Customizing Plots - -Seaborn plots are customized through function parameters (palette, hue, size, style) and by accessing the underlying Matplotlib axes after creation. `sns.set_theme()` and `sns.set_style()` change the global appearance. Seaborn's theming system makes it easy to produce clean, publication-ready charts with minimal code. \ No newline at end of file + +Seaborn plots are customized through function parameters (palette, hue, size, style) and by accessing the underlying Matplotlib axes after creation. `sns.set_theme()` and `sns.set_style()` change the global appearance. Seaborn's theming system makes it easy to produce clean, publication-ready charts with minimal code. + +Visit the following resources to learn more: + +- [@official@Controlling figure aesthetics](https://seaborn.pydata.org/tutorial/aesthetics.html) +- [@article@5 Ways to Transform Your Seaborn Data Visualisations](https://towardsdatascience.com/5-ways-to-transform-your-seaborn-data-visualisations-1ed2cb484e38/) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/dash@-dsafcJlA16H_ji0zQvKy.md b/src/data/roadmaps/python-data-analysis/content/dash@-dsafcJlA16H_ji0zQvKy.md index c90320cbd..4aca67389 100644 --- a/src/data/roadmaps/python-data-analysis/content/dash@-dsafcJlA16H_ji0zQvKy.md +++ b/src/data/roadmaps/python-data-analysis/content/dash@-dsafcJlA16H_ji0zQvKy.md @@ -1,3 +1,8 @@ # Dash - -Dash is a Python framework for building analytical web applications, developed by Plotly. It combines Plotly charts with reactive UI components and runs as a Flask web server. Dash provides more control and customization than Streamlit and is better suited for production-grade dashboards with complex interactivity. \ No newline at end of file + +Dash is a Python framework for building analytical web applications, developed by Plotly. It combines Plotly charts with reactive UI components and runs as a Flask web server. Dash provides more control and customization than Streamlit and is better suited for production-grade dashboards with complex interactivity. + +Visit the following resources to learn more: + +- [@official@Dash in 20 Minutes](https://dash.plotly.com/tutorial) +- [@video@Introduction to Dash Plotly - Data Visualization in Python](https://www.youtube.com/watch?v=hSPmj7mK6ng) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/dashboards@uyhlD_ZBCCd3IvrWZ5mlh.md b/src/data/roadmaps/python-data-analysis/content/dashboards@uyhlD_ZBCCd3IvrWZ5mlh.md index 9ae67cfeb..a8ec351b9 100644 --- a/src/data/roadmaps/python-data-analysis/content/dashboards@uyhlD_ZBCCd3IvrWZ5mlh.md +++ b/src/data/roadmaps/python-data-analysis/content/dashboards@uyhlD_ZBCCd3IvrWZ5mlh.md @@ -1,3 +1,8 @@ # Dashboards - -Dashboards combine multiple visualizations and controls into a single interface for monitoring and exploring data. Python provides several tools for building data dashboards that can be shared as web applications without requiring frontend development skills. The main options are Streamlit, Dash, and connecting to BI tools like Power BI and Tableau. \ No newline at end of file + +Dashboards combine multiple visualizations and controls into a single interface for monitoring and exploring data. Python provides several tools for building data dashboards that can be shared as web applications without requiring frontend development skills. The main options are Streamlit, Dash, and connecting to BI tools like Power BI and Tableau. + +Visit the following resources to learn more: + +- [@roadmap@Visit the Dedicated BI Analyst Roadmap](https://roadmap.sh/bi-analyst) +- [@article@What is a dashboard? A complete overview](https://www.tableau.com/dashboard/what-is-dashboard) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/dask@ApowKoA3C1B8oUr4-sltZ.md b/src/data/roadmaps/python-data-analysis/content/dask@ApowKoA3C1B8oUr4-sltZ.md index f4f225b30..3e63c23f0 100644 --- a/src/data/roadmaps/python-data-analysis/content/dask@ApowKoA3C1B8oUr4-sltZ.md +++ b/src/data/roadmaps/python-data-analysis/content/dask@ApowKoA3C1B8oUr4-sltZ.md @@ -1,3 +1,8 @@ # Dask - -Dask is a parallel computing library for Python that scales Pandas, NumPy, and Scikit-learn to larger-than-memory datasets. It breaks data into chunks and builds a task graph that is executed lazily. Dask DataFrames mirror the Pandas API, making it easy to adapt existing code for larger datasets without switching ecosystems. \ No newline at end of file + +Dask is a parallel computing library for Python that scales Pandas, NumPy, and Scikit-learn to larger-than-memory datasets. It breaks data into chunks and builds a task graph that is executed lazily. Dask DataFrames mirror the Pandas API, making it easy to adapt existing code for larger datasets without switching ecosystems. + +Visit the following resources to learn more: + +- [@official@Dask Tutorial](https://tutorial.dask.org/) +- [@video@Intro to Dask](https://www.youtube.com/watch?v=z18qjLu-Mw4&list=PLeDTMczuyDQ8S73cdc0PrnTO80kfzpgz2) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/data-cleaning@K53X-NB1S9XXSF6j8q1qu.md b/src/data/roadmaps/python-data-analysis/content/data-cleaning@K53X-NB1S9XXSF6j8q1qu.md index 4498c1842..e709f7484 100644 --- a/src/data/roadmaps/python-data-analysis/content/data-cleaning@K53X-NB1S9XXSF6j8q1qu.md +++ b/src/data/roadmaps/python-data-analysis/content/data-cleaning@K53X-NB1S9XXSF6j8q1qu.md @@ -1,3 +1,9 @@ # Data Cleaning - -Data cleaning identifies and resolves quality issues in raw data so it is accurate and consistent enough for analysis. In Python, cleaning is done primarily with Pandas and string processing tools. Common tasks include handling missing values, fixing data types, standardizing text, removing duplicates, and detecting outliers. \ No newline at end of file + +Data cleaning identifies and resolves quality issues in raw data so it is accurate and consistent enough for analysis. In Python, cleaning is done primarily with Pandas and string processing tools. Common tasks include handling missing values, fixing data types, standardizing text, removing duplicates, and detecting outliers. + +Visit the following resources to learn more: + +- [@article@Complete Guide to Data Cleaning in Python](https://www.dataquest.io/guide/data-cleaning-in-python-tutorial/) +- [@article@Pandas Data Cleaning](https://www.w3schools.com/python/pandas/pandas_cleaning.asp) +- [@video@Data Cleaning in Pandas | Python Pandas Tutorials](https://www.youtube.com/watch?v=bDhvCp3_lYw) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/data-pipelines@f3YL32h2gbU_G8kkNC3Vc.md b/src/data/roadmaps/python-data-analysis/content/data-pipelines@f3YL32h2gbU_G8kkNC3Vc.md index b8375630c..37a4e79db 100644 --- a/src/data/roadmaps/python-data-analysis/content/data-pipelines@f3YL32h2gbU_G8kkNC3Vc.md +++ b/src/data/roadmaps/python-data-analysis/content/data-pipelines@f3YL32h2gbU_G8kkNC3Vc.md @@ -1,3 +1,8 @@ # Data Pipelines - -Data pipelines automate the sequence of steps that move and transform data from sources to destinations. They encapsulate the full workflow β€” ingestion, cleaning, transformation, and output β€” as code. Orchestration tools like Airflow schedule and monitor these pipelines in production, ensuring they run reliably and their failures are caught and handled. \ No newline at end of file + +Data pipelines automate the sequence of steps that move and transform data from sources to destinations. They encapsulate the full workflow β€” ingestion, cleaning, transformation, and output β€” as code. Orchestration tools like Airflow schedule and monitor these pipelines in production, ensuring they run reliably and their failures are caught and handled. + +Visit the following resources to learn more: + +- [@article@What is a data pipeline?](https://www.ibm.com/think/topics/data-pipeline) +- [@video@Data Pipelines Explained](https://www.youtube.com/watch?v=6kEGUCrBEU0) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/defining-functions@0XtKHqPS6AfhQvhH5cu7Q.md b/src/data/roadmaps/python-data-analysis/content/defining-functions@0XtKHqPS6AfhQvhH5cu7Q.md index 8f0b663bb..0aad0dccc 100644 --- a/src/data/roadmaps/python-data-analysis/content/defining-functions@0XtKHqPS6AfhQvhH5cu7Q.md +++ b/src/data/roadmaps/python-data-analysis/content/defining-functions@0XtKHqPS6AfhQvhH5cu7Q.md @@ -1,3 +1,8 @@ # Defining Functions - -User-defined functions are created with the `def` keyword and encapsulate reusable logic. A function takes parameters, executes a body, and returns a value with `return`. Writing well-scoped functions makes analysis code modular, testable, and easier to apply across a dataset using Pandas' `apply()` method. \ No newline at end of file + +User-defined functions are created with the `def` keyword and encapsulate reusable logic. A function takes parameters, executes a body, and returns a value with `return`. Writing well-scoped functions makes analysis code modular, testable, and easier to apply across a dataset using Pandas' `apply()` method. + +Visit the following resources to learn more: + +- [@article@Python Return Multiple Values: 4 Methods & Examples](https://roadmap.sh/python/return-multiple-values) +- [@video@Python Functions - Visually Explained](https://www.youtube.com/watch?v=KW6qncswzHw) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/dictionaries@lReYQM-p8rrYcHkaqzbMr.md b/src/data/roadmaps/python-data-analysis/content/dictionaries@lReYQM-p8rrYcHkaqzbMr.md index 3160dd523..d724cf282 100644 --- a/src/data/roadmaps/python-data-analysis/content/dictionaries@lReYQM-p8rrYcHkaqzbMr.md +++ b/src/data/roadmaps/python-data-analysis/content/dictionaries@lReYQM-p8rrYcHkaqzbMr.md @@ -1,3 +1,9 @@ # Dictionaries - -Dictionaries store key-value pairs and provide fast lookup by key. They are used extensively in Python for mapping labels to values, building frequency counts, and configuring function arguments. \ No newline at end of file + +Dictionaries store key-value pairs and provide fast lookup by key. They are used extensively in Python for mapping labels to values, building frequency counts, and configuring function arguments. + +Visit the following resources to learn more: + +- [@official@Dictionaries](https://docs.python.org/3/tutorial/datastructures.html#dictionaries) +- [@article@Hashmaps in Python: Master Implementation and Use Cases](https://roadmap.sh/python/hashmap) +- [@video@Python dictionaries are easy πŸ“™](https://www.youtube.com/watch?v=MZZSMaEAC2g) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/distribution-plots@Xq7Ho2kWg5AtZK0G6z02U.md b/src/data/roadmaps/python-data-analysis/content/distribution-plots@Xq7Ho2kWg5AtZK0G6z02U.md index b72757e9c..de709aaf6 100644 --- a/src/data/roadmaps/python-data-analysis/content/distribution-plots@Xq7Ho2kWg5AtZK0G6z02U.md +++ b/src/data/roadmaps/python-data-analysis/content/distribution-plots@Xq7Ho2kWg5AtZK0G6z02U.md @@ -1,3 +1,7 @@ # Distribution plots - -Seaborn's distribution plots visualize the distribution of one or two variables. `sns.histplot()` and `sns.kdeplot()` show the shape of a single variable's distribution. `sns.displot()` combines both. `sns.pairplot()` shows pairwise distributions and relationships across all numeric columns in a DataFrame. \ No newline at end of file + +Seaborn's distribution plots visualize the distribution of one or two variables. `sns.histplot()` and `sns.kdeplot()` show the shape of a single variable's distribution. `sns.displot()` combines both. `sns.pairplot()` shows pairwise distributions and relationships across all numeric columns in a DataFrame. + +Visit the following resources to learn more: + +- [@official@Visualizing distributions of data](https://seaborn.pydata.org/tutorial/distributions.html) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/dropping-vs-imputing@xIS0cJ8ui0siBoZJr3-B_.md b/src/data/roadmaps/python-data-analysis/content/dropping-vs-imputing@xIS0cJ8ui0siBoZJr3-B_.md index 8c0797719..8b321b6cd 100644 --- a/src/data/roadmaps/python-data-analysis/content/dropping-vs-imputing@xIS0cJ8ui0siBoZJr3-B_.md +++ b/src/data/roadmaps/python-data-analysis/content/dropping-vs-imputing@xIS0cJ8ui0siBoZJr3-B_.md @@ -1,3 +1,9 @@ # Dropping vs. Imputing - -When handling missing values, dropping removes rows or columns with `dropna()`, while imputing fills them with a substitute value using `fillna()` or `SimpleImputer` from Scikit-learn. Dropping is appropriate when missing data is rare or random. Imputing is preferred when data is valuable or missing systematically, using the mean, median, mode, or a predicted value. \ No newline at end of file + +When handling missing values, dropping removes rows or columns with `dropna()`, while imputing fills them with a substitute value using `fillna()` or `SimpleImputer` from Scikit-learn. Dropping is appropriate when missing data is rare or random. Imputing is preferred when data is valuable or missing systematically, using the mean, median, mode, or a predicted value. + +Visit the following resources to learn more: + +- [@official@pandas.DataFrame.fillna](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.fillna.html) +- [@article@Pandas DataFrame fillna() Method](https://www.w3schools.com/python/pandas/ref_df_fillna.asp) +- [@video@Python Pandas Tutorial 5: Handle Missing Data: fillna, dropna, interpolate](https://www.youtube.com/watch?v=EaGbS7eWSs0) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/duckdb@a6HiqhsJJijqObw9pNOqb.md b/src/data/roadmaps/python-data-analysis/content/duckdb@a6HiqhsJJijqObw9pNOqb.md index 4abc490c9..364aad49b 100644 --- a/src/data/roadmaps/python-data-analysis/content/duckdb@a6HiqhsJJijqObw9pNOqb.md +++ b/src/data/roadmaps/python-data-analysis/content/duckdb@a6HiqhsJJijqObw9pNOqb.md @@ -1,3 +1,9 @@ # DuckDB - -DuckDB is an in-process analytical database designed for fast SQL queries on large datasets stored as files or in memory. It can query CSV, Parquet, and Pandas DataFrames directly with SQL syntax. DuckDB is increasingly used in data analysis workflows as a fast alternative to loading data into a full database system. \ No newline at end of file + +DuckDB is an in-process analytical database designed for fast SQL queries on large datasets stored as files or in memory. It can query CSV, Parquet, and Pandas DataFrames directly with SQL syntax. DuckDB is increasingly used in data analysis workflows as a fast alternative to loading data into a full database system. + +Visit the following resources to learn more: + +- [@official@DuckDB Docs](https://duckdb.org/docs/current/clients/python/overview) +- [@article@Introducing DuckDB](https://realpython.com/python-duckdb/) +- [@video@Try DuckDB for SQL on Pandas](https://www.youtube.com/watch?v=8SYQtpSk_OI) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/encoding-categories@1bHv3_acNts_BeOn3l6DF.md b/src/data/roadmaps/python-data-analysis/content/encoding-categories@1bHv3_acNts_BeOn3l6DF.md index 2aa9505c8..cce4128d5 100644 --- a/src/data/roadmaps/python-data-analysis/content/encoding-categories@1bHv3_acNts_BeOn3l6DF.md +++ b/src/data/roadmaps/python-data-analysis/content/encoding-categories@1bHv3_acNts_BeOn3l6DF.md @@ -1,3 +1,8 @@ # Encoding Categories - -Categorical encoding converts text category labels into numerical values that machine learning algorithms can process. Common approaches include label encoding (assigning each category an integer), one-hot encoding (creating binary columns for each category with `pd.get_dummies()`), and ordinal encoding for ordered categories. \ No newline at end of file + +Categorical encoding converts text category labels into numerical values that machine learning algorithms can process. Common approaches include label encoding (assigning each category an integer), one-hot encoding (creating binary columns for each category with `pd.get_dummies()`), and ordinal encoding for ordered categories. + +Visit the following resources to learn more: + +- [@official@Categorical data](https://pandas.pydata.org/docs/user_guide/categorical.html) +- [@article@Encoding Categorical Variables: One-hot vs Dummy Encoding](https://towardsdatascience.com/encoding-categorical-variables-one-hot-vs-dummy-encoding-6d5b9c46e2db/) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/environment-setup@8Z20apU8ugXwdH8rRZAcr.md b/src/data/roadmaps/python-data-analysis/content/environment-setup@8Z20apU8ugXwdH8rRZAcr.md index f1109ca56..b05b5ebb9 100644 --- a/src/data/roadmaps/python-data-analysis/content/environment-setup@8Z20apU8ugXwdH8rRZAcr.md +++ b/src/data/roadmaps/python-data-analysis/content/environment-setup@8Z20apU8ugXwdH8rRZAcr.md @@ -1,3 +1,7 @@ # Environment Setup - -Setting up a proper Python environment for data analysis involves choosing a package manager, managing dependencies, and selecting a development environment. A well-configured environment ensures reproducibility and avoids package conflicts. The main tools are pip and conda for package management, and virtual environments for isolation. \ No newline at end of file + +Setting up a proper Python environment for data analysis involves choosing a package manager, managing dependencies, and selecting a development environment. A well-configured environment ensures reproducibility and avoids package conflicts. The main tools are pip and conda for package management, and virtual environments for isolation. + +Visit the following resources to learn more: + +- [@video@Setting Up A Python Environment for Data Analysis and Machine Learning](https://www.youtube.com/watch?v=NDFMa5FSQuI) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/excel@OwEBYI4u7nb9-UGOR9SQf.md b/src/data/roadmaps/python-data-analysis/content/excel@OwEBYI4u7nb9-UGOR9SQf.md index 46c8dd177..0768cea01 100644 --- a/src/data/roadmaps/python-data-analysis/content/excel@OwEBYI4u7nb9-UGOR9SQf.md +++ b/src/data/roadmaps/python-data-analysis/content/excel@OwEBYI4u7nb9-UGOR9SQf.md @@ -1,3 +1,8 @@ # Excel - -Excel files (`.xlsx`, `.xls`) are loaded with `pd.read_excel()`, which supports selecting sheets, skipping rows, and reading specific columns. The `openpyxl` library is required for `.xlsx` files. Excel is common in business environments, and analysts frequently need to read and write it as part of reporting workflows. \ No newline at end of file + +Excel files (`.xlsx`, `.xls`) are loaded with `pd.read_excel()`, which supports selecting sheets, skipping rows, and reading specific columns. The `openpyxl` library is required for `.xlsx` files. Excel is common in business environments, and analysts frequently need to read and write it as part of reporting workflows. + +Visit the following resources to learn more: + +- [@official@pandas.read_excel](https://pandas.pydata.org/docs/reference/api/pandas.read_excel.html) +- [@article@Pandas read_excel: Reading Excel Files in Python](https://www.digitalocean.com/community/tutorials/pandas-read_excel-reading-excel-file-in-python) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/exploratory-data-analysis@eObhSoXcbqLdwegzaWmi7.md b/src/data/roadmaps/python-data-analysis/content/exploratory-data-analysis@eObhSoXcbqLdwegzaWmi7.md index 35dc69499..c14c561ad 100644 --- a/src/data/roadmaps/python-data-analysis/content/exploratory-data-analysis@eObhSoXcbqLdwegzaWmi7.md +++ b/src/data/roadmaps/python-data-analysis/content/exploratory-data-analysis@eObhSoXcbqLdwegzaWmi7.md @@ -1,3 +1,9 @@ # Exploratory Data Analysis - -Exploratory Data Analysis (EDA) is the process of examining a dataset to understand its structure, distributions, and relationships before formal modeling. It combines descriptive statistics and visualizations to surface patterns, anomalies, and hypotheses. EDA guides subsequent cleaning decisions and model choices. \ No newline at end of file + +Exploratory Data Analysis (EDA) is the process of examining a dataset to understand its structure, distributions, and relationships before formal modeling. It combines descriptive statistics and visualizations to surface patterns, anomalies, and hypotheses. EDA guides subsequent cleaning decisions and model choices. + +Visit the following resources to learn more: + +- [@article@Exploratory Statistical Data Analysis with a Real Dataset using Pandas](https://medium.com/data-science/exploratory-statistical-data-analysis-with-a-real-dataset-using-pandas-208007798b92) +- [@article@Pandas Profiling – Easy Exploratory Data Analysis in Python](https://towardsdatascience.com/pandas-profiling-easy-exploratory-data-analysis-in-python-65d6d0e23650/) +- [@video@Exploratory Data Analysis in Pandas | Python Pandas Tutorials](https://www.youtube.com/watch?v=Liv6eeb1VfE) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/filtering--querying@wxO7f8wh6JTXHYDiPcbmy.md b/src/data/roadmaps/python-data-analysis/content/filtering--querying@wxO7f8wh6JTXHYDiPcbmy.md index adea10d98..c2f108d5e 100644 --- a/src/data/roadmaps/python-data-analysis/content/filtering--querying@wxO7f8wh6JTXHYDiPcbmy.md +++ b/src/data/roadmaps/python-data-analysis/content/filtering--querying@wxO7f8wh6JTXHYDiPcbmy.md @@ -1,3 +1,10 @@ # Filtering & Querying - -Filtering in Pandas selects rows that meet specified conditions. Boolean masks, the `.query()` method, and `.isin()` are common approaches. Multiple conditions can be combined with `&` (and) and `|` (or), and the `.query()` method allows SQL-like string syntax for readable filtering expressions. \ No newline at end of file + +Filtering in Pandas selects rows that meet specified conditions. Boolean masks, the `.query()` method, and `.isin()` are common approaches. Multiple conditions can be combined with `&` (and) and `|` (or), and the `.query()` method allows SQL-like string syntax for readable filtering expressions. + +Visit the following resources to learn more: + +- [@official@pandas.DataFrame.query](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.query.htmlme.query) +- [@official@How do I select a subset of a DataFrame?](https://pandas.pydata.org/docs/getting_started/intro_tutorials/03_subset_data.html) +- [@article@10 Elegant Ways to Filter Pandas DataFrames](https://towardsdatascience.com/stop-writing-messy-boolean-masks-10-elegant-ways-to-filter-pandas-dataframes/) +- [@video@Filtering Columns and Rows in Pandas](https://www.youtube.com/watch?v=kB7FV-ijdqE) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/floats@aLIXcqB258WNrdASuN_7G.md b/src/data/roadmaps/python-data-analysis/content/floats@aLIXcqB258WNrdASuN_7G.md index 5ed6951a2..0e16338ce 100644 --- a/src/data/roadmaps/python-data-analysis/content/floats@aLIXcqB258WNrdASuN_7G.md +++ b/src/data/roadmaps/python-data-analysis/content/floats@aLIXcqB258WNrdASuN_7G.md @@ -1,3 +1,8 @@ # Floats - -Floats (`float`) represent real numbers with a decimal point. Most numerical data in analysis involves floats, including prices, measurements, and probabilities. Floating-point arithmetic has precision limitations that can cause small rounding errors, which are important to be aware of in financial and scientific calculations. \ No newline at end of file + +Floats (`float`) represent real numbers with a decimal point. Most numerical data in analysis involves floats, including prices, measurements, and probabilities. Floating-point arithmetic has precision limitations that can cause small rounding errors, which are important to be aware of in financial and scientific calculations. + +Visit the following resources to learn more: + +- [@official@Built-in Types](https://docs.python.org/3/library/stdtypes.html) +- [@article@float](https://realpython.com/ref/builtin-types/float/) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/forward--backward-fill@9mFT0z5yT446_-nGdG492.md b/src/data/roadmaps/python-data-analysis/content/forward--backward-fill@9mFT0z5yT446_-nGdG492.md index fcadd9813..145846673 100644 --- a/src/data/roadmaps/python-data-analysis/content/forward--backward-fill@9mFT0z5yT446_-nGdG492.md +++ b/src/data/roadmaps/python-data-analysis/content/forward--backward-fill@9mFT0z5yT446_-nGdG492.md @@ -1,3 +1,9 @@ # Forward / Backward Fill - -Forward fill (`ffill`) propagates the last valid value forward to fill subsequent missing entries. Backward fill (`bfill`) does the reverse, filling from the next valid value. Both are commonly used for time series data where missing values represent periods where the previous or next observation is the best estimate. \ No newline at end of file + +Forward fill (`ffill`) propagates the last valid value forward to fill subsequent missing entries. Backward fill (`bfill`) does the reverse, filling from the next valid value. Both are commonly used for time series data where missing values represent periods where the previous or next observation is the best estimate. + +Visit the following resources to learn more: + +- [@official@pandas.DataFrame.ffill](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.ffill.html) +- [@official@pandas.DataFrame.bfill](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.bfill.html) +- [@article@How to Fill Missing Data with Pandas](https://towardsdatascience.com/how-to-fill-missing-data-with-pandas-8cb875362a0d/) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/functions--methods@kyrS5ez7_85AnNt8Bojz_.md b/src/data/roadmaps/python-data-analysis/content/functions--methods@kyrS5ez7_85AnNt8Bojz_.md index ccc29aae6..644d6c747 100644 --- a/src/data/roadmaps/python-data-analysis/content/functions--methods@kyrS5ez7_85AnNt8Bojz_.md +++ b/src/data/roadmaps/python-data-analysis/content/functions--methods@kyrS5ez7_85AnNt8Bojz_.md @@ -1,3 +1,9 @@ # Functions & Methods - -Functions are reusable blocks of code that take inputs, perform operations, and return outputs. They encapsulate cleaning steps, transformations, and calculations that need to be applied consistently. Python supports built-in functions, user-defined functions, and anonymous lambda functions. \ No newline at end of file + +Functions are reusable blocks of code that take inputs, perform operations, and return outputs. They encapsulate cleaning steps, transformations, and calculations that need to be applied consistently. Python supports built-in functions, user-defined functions, and anonymous lambda functions. + +Visit the following resources to learn more: + +- [@article@Python Functions](https://www.w3schools.com/python/python_functions.asp) +- [@article@Python Methods, Functions, & Libraries](https://mode.com/python-tutorial/python-methods-functions-and-libraries) +- [@video@Functions in Python are easy](https://www.youtube.com/watch?v=89cGQjB5R4M) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/google-colab@c8uLNJMr6UZfvjFwFhxvd.md b/src/data/roadmaps/python-data-analysis/content/google-colab@c8uLNJMr6UZfvjFwFhxvd.md index fc0083829..9b7411869 100644 --- a/src/data/roadmaps/python-data-analysis/content/google-colab@c8uLNJMr6UZfvjFwFhxvd.md +++ b/src/data/roadmaps/python-data-analysis/content/google-colab@c8uLNJMr6UZfvjFwFhxvd.md @@ -1,3 +1,9 @@ # Google Colab - -Google Colab is a cloud-hosted Jupyter notebook environment from Google. It requires no local setup and provides free access to GPUs and TPUs, making it popular for machine learning work. Colab notebooks are stored in Google Drive and can be shared like any other document. \ No newline at end of file + +Google Colab is a cloud-hosted Jupyter notebook environment from Google. It requires no local setup and provides free access to GPUs and TPUs, making it popular for machine learning work. Colab notebooks are stored in Google Drive and can be shared like any other document. + +Visit the following resources to learn more: + +- [@official@Google Colab](https://colab.research.google.com/) +- [@official@Python Basics in Colab](https://colab.research.google.com/github/data-psl/lectures2020/blob/master/notebooks/01_python_basics.ipynb) +- [@video@Google Colab Tutorial for Beginners | Get Started with Google Colab](https://www.youtube.com/watch?v=RLYoEyIHL6A) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/groupby--aggregation@KnSv5dPToipoz2RF1dDnO.md b/src/data/roadmaps/python-data-analysis/content/groupby--aggregation@KnSv5dPToipoz2RF1dDnO.md index 66c0064d7..f425904b6 100644 --- a/src/data/roadmaps/python-data-analysis/content/groupby--aggregation@KnSv5dPToipoz2RF1dDnO.md +++ b/src/data/roadmaps/python-data-analysis/content/groupby--aggregation@KnSv5dPToipoz2RF1dDnO.md @@ -1,3 +1,8 @@ # Groupby & Aggregation - -`groupby()` splits a DataFrame into groups based on one or more columns, applies a function to each group, and combines the results. Common aggregation functions include `sum()`, `mean()`, `count()`, `min()`, `max()`, and custom functions via `agg()`. This split-apply-combine pattern is one of the most powerful features of Pandas. \ No newline at end of file + +`groupby()` splits a DataFrame into groups based on one or more columns, applies a function to each group, and combines the results. Common aggregation functions include `sum()`, `mean()`, `count()`, `min()`, `max()`, and custom functions via `agg()`. This split-apply-combine pattern is one of the most powerful features of Pandas. + +Visit the following resources to learn more: + +- [@official@pandas.DataFrame.groupby](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.groupby.html) +- [@video@Group By and Aggregate Functions in Pandas |](https://www.youtube.com/watch?v=VRmXto2YA2I) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/heatmaps@Bk3z1CHnrnV6ii2Y0BF8G.md b/src/data/roadmaps/python-data-analysis/content/heatmaps@Bk3z1CHnrnV6ii2Y0BF8G.md index 4c3ceeae9..2adb2cf0b 100644 --- a/src/data/roadmaps/python-data-analysis/content/heatmaps@Bk3z1CHnrnV6ii2Y0BF8G.md +++ b/src/data/roadmaps/python-data-analysis/content/heatmaps@Bk3z1CHnrnV6ii2Y0BF8G.md @@ -1,3 +1,9 @@ # Heatmaps - -`sns.heatmap()` visualizes matrix-style data using color intensity. It is most commonly used to display correlation matrices and pivot tables. Color maps, annotations, and masking options allow the heatmap to be customized for readability. Heatmaps are an effective way to show patterns across two categorical dimensions. \ No newline at end of file + +`sns.heatmap()` visualizes matrix-style data using color intensity. It is most commonly used to display correlation matrices and pivot tables. Color maps, annotations, and masking options allow the heatmap to be customized for readability. Heatmaps are an effective way to show patterns across two categorical dimensions. + +Visit the following resources to learn more: + +- [@official@seaborn.heatmap](https://seaborn.pydata.org/generated/seaborn.heatmap.html) +- [@article@Data Visualization with Seaborn: Heatmaps](https://medium.com/@1zeyneper/data-visualization-with-seaborn-heatmaps-58abfadd79d5) +- [@video@Seaborn heatmap | How to make a heatmap in Python Seaborn and adjust the heatmap style](https://www.youtube.com/watch?v=0U9cs2V-Mqc) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/histogram@oRaag-FOl50NJM0y4joBm.md b/src/data/roadmaps/python-data-analysis/content/histogram@oRaag-FOl50NJM0y4joBm.md index dac696609..9a46018db 100644 --- a/src/data/roadmaps/python-data-analysis/content/histogram@oRaag-FOl50NJM0y4joBm.md +++ b/src/data/roadmaps/python-data-analysis/content/histogram@oRaag-FOl50NJM0y4joBm.md @@ -1,3 +1,9 @@ # Histogram - -A histogram groups numeric values into bins and shows the count or frequency of each bin as a bar. It is the primary tool for visualizing the distribution of a single variable: its shape, center, spread, and whether it is skewed or has multiple peaks. `df['col'].hist()` and Matplotlib's `plt.hist()` are the standard ways to create one. \ No newline at end of file + +A histogram groups numeric values into bins and shows the count or frequency of each bin as a bar. It is the primary tool for visualizing the distribution of a single variable: its shape, center, spread, and whether it is skewed or has multiple peaks. `df['col'].hist()` and Matplotlib's `plt.hist()` are the standard ways to create one. + +Visit the following resources to learn more: + +- [@official@matplotlib.pyplot.hist](https://matplotlib.org/stable/api/_as_gen/matplotlib.pyplot.hist.html) +- [@article@Matplotlib Histograms](https://www.w3schools.com/PYTHON/matplotlib_histograms.asp) +- [@video@Matplotlib histograms in 6 minutes! πŸ””](https://www.youtube.com/watch?v=2E6fDoz7LuU) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/indexing--slicing@PlFcaQsvFSK0Y-LKGu0eH.md b/src/data/roadmaps/python-data-analysis/content/indexing--slicing@PlFcaQsvFSK0Y-LKGu0eH.md index 0be9c4b98..73c257116 100644 --- a/src/data/roadmaps/python-data-analysis/content/indexing--slicing@PlFcaQsvFSK0Y-LKGu0eH.md +++ b/src/data/roadmaps/python-data-analysis/content/indexing--slicing@PlFcaQsvFSK0Y-LKGu0eH.md @@ -1,3 +1,10 @@ # Indexing & Slicing - -Pandas provides two primary indexing systems: `.loc[]` for label-based selection and `.iloc[]` for position-based selection. Both work on rows, columns, or both simultaneously. Boolean indexing with a condition (e.g., `df[df['age'] > 30]`) is the most common way to filter rows. \ No newline at end of file + +Pandas provides two primary indexing systems: `.loc[]` for label-based selection and `.iloc[]` for position-based selection. Both work on rows, columns, or both simultaneously. Boolean indexing with a condition (e.g., `df[df['age'] > 30]`) is the most common way to filter rows. + +Visit the following resources to learn more: + +- [@official@pandas.DataFrame.loc](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.loc.html) +- [@official@pandas.DataFrame.iloc](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.iloc.html) +- [@article@Iloc vs Loc in Pandas: A Guide with Examples](https://www.analyticsvidhya.com/blog/2026/03/iloc-vs-loc-in-pandas/) +- [@video@Pandas loc and iloc](https://www.youtube.com/watch?v=naRQyRZrXCE) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/indexing--slicing@mAOljEWt_0Hb9jLBKuc-H.md b/src/data/roadmaps/python-data-analysis/content/indexing--slicing@mAOljEWt_0Hb9jLBKuc-H.md index 67d5bd55f..eaf56b101 100644 --- a/src/data/roadmaps/python-data-analysis/content/indexing--slicing@mAOljEWt_0Hb9jLBKuc-H.md +++ b/src/data/roadmaps/python-data-analysis/content/indexing--slicing@mAOljEWt_0Hb9jLBKuc-H.md @@ -1,3 +1,9 @@ # Array Operations - -NumPy supports a wide range of array operations: element-wise arithmetic, aggregation functions (`sum`, `mean`, `std`, `min`, `max`), reshaping, stacking, and splitting. These operations are vectorized, meaning they apply to the entire array at once without explicit loops, making them highly efficient. \ No newline at end of file + +NumPy supports a wide range of array operations: element-wise arithmetic, aggregation functions (`sum`, `mean`, `std`, `min`, `max`), reshaping, stacking, and splitting. These operations are vectorized, meaning they apply to the entire array at once without explicit loops, making them highly efficient. + +Visit the following resources to learn more: + +- [@official@Indexing on ndarrays](https://numpy.org/doc/stable/user/basics.indexing.html) +- [@article@NumPy Array Indexing](https://www.w3schools.com/python/numpy/numpy_array_indexing.asp) +- [@article@NumPy Array Slicing](https://www.w3schools.com/python/numpy/numpy_array_slicing.asp) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/integers@_-bUq5jtewDmsmnzYw1-q.md b/src/data/roadmaps/python-data-analysis/content/integers@_-bUq5jtewDmsmnzYw1-q.md index 2e79167f2..f6e28598a 100644 --- a/src/data/roadmaps/python-data-analysis/content/integers@_-bUq5jtewDmsmnzYw1-q.md +++ b/src/data/roadmaps/python-data-analysis/content/integers@_-bUq5jtewDmsmnzYw1-q.md @@ -1,3 +1,8 @@ # Integers - -Integers (`int`) are whole numbers without a decimal point. They appear as counts, indices, IDs, and categorical encodings. Python integers have arbitrary precision, meaning they do not overflow like integers in lower-level languages. \ No newline at end of file + +Integers (`int`) are whole numbers without a decimal point. They appear as counts, indices, IDs, and categorical encodings. Python integers have arbitrary precision, meaning they do not overflow like integers in lower-level languages. + +Visit the following resources to learn more: + +- [@official@Built-in Types](https://docs.python.org/3/library/stdtypes.html) +- [@article@int](https://realpython.com/ref/builtin-types/int/) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/interactive-visualization@ES8zVixZ5wZ0tW6beOWxJ.md b/src/data/roadmaps/python-data-analysis/content/interactive-visualization@ES8zVixZ5wZ0tW6beOWxJ.md index 22b4db6f6..1e0dbc73c 100644 --- a/src/data/roadmaps/python-data-analysis/content/interactive-visualization@ES8zVixZ5wZ0tW6beOWxJ.md +++ b/src/data/roadmaps/python-data-analysis/content/interactive-visualization@ES8zVixZ5wZ0tW6beOWxJ.md @@ -1,3 +1,7 @@ # Interactive Visualization - -Interactive visualizations allow users to explore data by hovering, zooming, panning, and filtering directly in the chart. They are more engaging than static plots for dashboards and reports where the audience needs to examine specific data points. Python's main interactive visualization libraries are Plotly and Altair. \ No newline at end of file + +Interactive visualizations allow users to explore data by hovering, zooming, panning, and filtering directly in the chart. They are more engaging than static plots for dashboards and reports where the audience needs to examine specific data points. Python's main interactive visualization libraries are Plotly and Altair. + +Visit the following resources to learn more: + +- [@article@Top 10 Python Data Visualization Libraries](https://reflex.dev/blog/top-10-data-visualization-libraries/) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/introduction@GISOFMKvnBys0O0IMpz2J.md b/src/data/roadmaps/python-data-analysis/content/introduction@GISOFMKvnBys0O0IMpz2J.md index e00ebb33f..129c65343 100644 --- a/src/data/roadmaps/python-data-analysis/content/introduction@GISOFMKvnBys0O0IMpz2J.md +++ b/src/data/roadmaps/python-data-analysis/content/introduction@GISOFMKvnBys0O0IMpz2J.md @@ -1,3 +1,10 @@ # Introduction - -Python is the dominant language for data analysis due to its readable syntax, rich ecosystem of libraries, and strong community support. Getting started requires understanding the core language features: operators, data types, control flow, and data structures. These fundamentals apply directly to every data manipulation and analysis task that follows. \ No newline at end of file + +Python is the dominant language for data analysis due to its readable syntax, rich ecosystem of libraries, and strong community support. Getting started requires understanding the core language features: operators, data types, control flow, and data structures. These fundamentals apply directly to every data manipulation and analysis task that follows. + +Visit the following resources to learn more: + +- [@book@Python for Data Analysis](https://www.lkhibra.ma/books/Python-for-Data-Analysis.pdf) +- [@article@What Does a Data Analyst Do?](https://roadmap.sh/data-analyst/what-does-a-data-analyst-do) +- [@article@How Long Does It Really Take To Learn Python? My Experience](https://roadmap.sh/python/how-long-does-it-take-to-learn) +- [@video@Python for Data Analytics - Full Course for Beginners](https://www.youtube.com/watch?v=wUSDVGivd-8) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/iqr@ucEFYVtXIvuYKjudK3pwB.md b/src/data/roadmaps/python-data-analysis/content/iqr@ucEFYVtXIvuYKjudK3pwB.md index 5f031bf6c..9f7cb9e1d 100644 --- a/src/data/roadmaps/python-data-analysis/content/iqr@ucEFYVtXIvuYKjudK3pwB.md +++ b/src/data/roadmaps/python-data-analysis/content/iqr@ucEFYVtXIvuYKjudK3pwB.md @@ -1,3 +1,9 @@ # IQR - -The Interquartile Range (IQR) is the difference between the 75th percentile (Q3) and 25th percentile (Q1) of a dataset. Outliers are commonly defined as values below Q1 βˆ’ 1.5Γ—IQR or above Q3 + 1.5Γ—IQR. The IQR method is robust to extreme values and is the basis for the box plot's whiskers. \ No newline at end of file + +The Interquartile Range (IQR) is the difference between the 75th percentile (Q3) and 25th percentile (Q1) of a dataset. Outliers are commonly defined as values below Q1 βˆ’ 1.5Γ—IQR or above Q3 + 1.5Γ—IQR. The IQR method is robust to extreme values and is the basis for the box plot's whiskers. + +Visit the following resources to learn more: + +- [@article@Using Pandas IQR: 3 Essential Steps](https://medium.com/@heyamit10/using-pandas-iqr-3-essential-steps-f5cf73a390ba) +- [@article@How to detect outliers using IQR and Boxplots?](https://machinelearningplus.com/machine-learning/how-to-detect-outliers-using-iqr-and-boxplots/) +- [@video@Outlier detection and removal using IQR](https://www.youtube.com/watch?v=A3gClkblXK8) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/isnull-isna@bcODvfBqCUwLVXa8HKOFB.md b/src/data/roadmaps/python-data-analysis/content/isnull-isna@bcODvfBqCUwLVXa8HKOFB.md index 044ac4272..840be5545 100644 --- a/src/data/roadmaps/python-data-analysis/content/isnull-isna@bcODvfBqCUwLVXa8HKOFB.md +++ b/src/data/roadmaps/python-data-analysis/content/isnull-isna@bcODvfBqCUwLVXa8HKOFB.md @@ -1,3 +1,10 @@ # isnull, isna - -`isnull()` and `isna()` are equivalent Pandas methods that return a boolean DataFrame or Series indicating which values are missing (NaN). They are the first step in assessing data completeness. Combined with `.sum()`, they give a count of missing values per column, and with boolean indexing they select rows with missing data. \ No newline at end of file + +`isnull()` and `isna()` are equivalent Pandas methods that return a boolean DataFrame or Series indicating which values are missing (NaN). They are the first step in assessing data completeness. Combined with `.sum()`, they give a count of missing values per column, and with boolean indexing they select rows with missing data. + +Visit the following resources to learn more: + +- [@official@Working with missing data](https://pandas.pydata.org/docs/user_guide/missing_data.html) +- [@official@pandas.isnull](https://pandas.pydata.org/docs/reference/api/pandas.isnull.html) +- [@official@pandas.isna](https://pandas.pydata.org/docs/reference/api/pandas.isna.html) +- [@article@Handling Missing Values with Pandas](https://towardsdatascience.com/handling-missing-values-with-pandas-b876bf6f008f/) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/json@VEetEvKGNzAO9QdZY2g86.md b/src/data/roadmaps/python-data-analysis/content/json@VEetEvKGNzAO9QdZY2g86.md index 6e88871a2..708fc0835 100644 --- a/src/data/roadmaps/python-data-analysis/content/json@VEetEvKGNzAO9QdZY2g86.md +++ b/src/data/roadmaps/python-data-analysis/content/json@VEetEvKGNzAO9QdZY2g86.md @@ -1,3 +1,8 @@ # JSON - -JSON (JavaScript Object Notation) is a text format for structured data commonly returned by APIs. `pd.read_json()` converts JSON into a DataFrame, though nested structures often require normalization with `pd.json_normalize()`. JSON is flexible but can be irregular in structure, requiring careful handling of missing fields. \ No newline at end of file + +JSON (JavaScript Object Notation) is a text format for structured data commonly returned by APIs. `pd.read_json()` converts JSON into a DataFrame, though nested structures often require normalization with `pd.json_normalize()`. JSON is flexible but can be irregular in structure, requiring careful handling of missing fields. + +Visit the following resources to learn more: + +- [@official@pandas.read_json](https://pandas.pydata.org/docs/reference/api/pandas.read_json.html) +- [@article@Pandas Read JSON](https://www.w3schools.com/python/pandas/pandas_json.asp) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/jupyterlab@2m8uQRGaVZUd4mfm7bRou.md b/src/data/roadmaps/python-data-analysis/content/jupyterlab@2m8uQRGaVZUd4mfm7bRou.md index 3837579bf..4af06b4b9 100644 --- a/src/data/roadmaps/python-data-analysis/content/jupyterlab@2m8uQRGaVZUd4mfm7bRou.md +++ b/src/data/roadmaps/python-data-analysis/content/jupyterlab@2m8uQRGaVZUd4mfm7bRou.md @@ -1,3 +1,8 @@ # JupyterLab - -JupyterLab is the modern, full-featured successor to the classic Jupyter Notebook interface. It supports notebooks where code, output, and narrative text coexist in a single document, while adding a tabbed layout, a file browser, a terminal, and support for multiple file types side by side. It is the standard environment for exploratory data analysis because results are visible immediately after each cell is run. \ No newline at end of file + +JupyterLab is the modern, full-featured successor to the classic Jupyter Notebook interface. It supports notebooks where code, output, and narrative text coexist in a single document, while adding a tabbed layout, a file browser, a terminal, and support for multiple file types side by side. It is the standard environment for exploratory data analysis because results are visible immediately after each cell is run. + +Visit the following resources to learn more: + +- [@official@Get Started](https://jupyterlab.readthedocs.io/en/stable/getting_started/overview.html) +- [@video@Jupyter Notebook Complete Beginner Guide](https://www.youtube.com/watch?v=5pf0_bpNbkw) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/lambda-functions@AzsZhpK-wuWtafVFJMhdY.md b/src/data/roadmaps/python-data-analysis/content/lambda-functions@AzsZhpK-wuWtafVFJMhdY.md index e69bd9360..e94054a13 100644 --- a/src/data/roadmaps/python-data-analysis/content/lambda-functions@AzsZhpK-wuWtafVFJMhdY.md +++ b/src/data/roadmaps/python-data-analysis/content/lambda-functions@AzsZhpK-wuWtafVFJMhdY.md @@ -1,3 +1,8 @@ # Lambda Functions - -Lambda functions are anonymous, single-expression functions defined with the `lambda` keyword. They are used for short, throwaway operations, particularly as arguments to functions like `map()`, `filter()`, and Pandas' `apply()`. For example: `df['col'].apply(lambda x: x * 2)`. \ No newline at end of file + +Lambda functions are anonymous, single-expression functions defined with the `lambda` keyword. They are used for short, throwaway operations, particularly as arguments to functions like `map()`, `filter()`, and Pandas' `apply()`. For example: `df['col'].apply(lambda x: x * 2)`. + +Visit the following resources to learn more: + +- [@article@Python Lambda](https://www.w3schools.com/python/python_lambda.asp) +- [@video@Python Lambda Functions Explained](https://www.youtube.com/watch?v=HQNiSfb795A) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/linear-algebra-basics@xHDrwKDE_mrSihCsPQMBF.md b/src/data/roadmaps/python-data-analysis/content/linear-algebra-basics@xHDrwKDE_mrSihCsPQMBF.md index 875e42301..a542bb5d1 100644 --- a/src/data/roadmaps/python-data-analysis/content/linear-algebra-basics@xHDrwKDE_mrSihCsPQMBF.md +++ b/src/data/roadmaps/python-data-analysis/content/linear-algebra-basics@xHDrwKDE_mrSihCsPQMBF.md @@ -1,3 +1,8 @@ # Linear Algebra Basics - -NumPy provides linear algebra operations, including matrix multiplication (`np.dot`, `@`), matrix inversion, determinants, and eigenvalues. These are used in statistics (covariance matrices), machine learning (feature transformations), and scientific computing. Understanding the basics of matrix operations is useful for reading ML algorithm implementations. \ No newline at end of file + +NumPy provides linear algebra operations, including matrix multiplication (`np.dot`, `@`), matrix inversion, determinants, and eigenvalues. These are used in statistics (covariance matrices), machine learning (feature transformations), and scientific computing. Understanding the basics of matrix operations is useful for reading ML algorithm implementations. + +Visit the following resources to learn more: + +- [@article@Numpy Linear Algebra](https://www.programiz.com/python-programming/numpy/linear-algebra) +- [@video@Learn NumPy in 1 hour! πŸ”’](https://www.youtube.com/watch?v=VXU4LSAQDSc) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/list-comprehensions@la9oiXURvw6AH-Vuh-WqT.md b/src/data/roadmaps/python-data-analysis/content/list-comprehensions@la9oiXURvw6AH-Vuh-WqT.md index 244d8bab1..6ed18352a 100644 --- a/src/data/roadmaps/python-data-analysis/content/list-comprehensions@la9oiXURvw6AH-Vuh-WqT.md +++ b/src/data/roadmaps/python-data-analysis/content/list-comprehensions@la9oiXURvw6AH-Vuh-WqT.md @@ -1,3 +1,8 @@ # List Comprehensions - -List comprehensions provide a concise syntax for creating lists by applying an expression to each item in an iterable, optionally filtering with a condition. For example: `[x**2 for x in range(10) if x % 2 == 0]`. They are faster and more readable than equivalent `for` loops for simple transformations. \ No newline at end of file + +List comprehensions provide a concise syntax for creating lists by applying an expression to each item in an iterable, optionally filtering with a condition. For example: `[x**2 for x in range(10) if x % 2 == 0]`. They are faster and more readable than equivalent `for` loops for simple transformations. + +Visit the following resources to learn more: + +- [@official@List Comprehensions](https://docs.python.org/3/tutorial/datastructures.html#list-comprehensions) +- [@article@When to Use a List Comprehension in Python Quiz](https://realpython.com/quizzes/list-comprehension-python/) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/lists@F4dp3Ip8G-EI1hprtv3do.md b/src/data/roadmaps/python-data-analysis/content/lists@F4dp3Ip8G-EI1hprtv3do.md index b9660e005..72b8fd50c 100644 --- a/src/data/roadmaps/python-data-analysis/content/lists@F4dp3Ip8G-EI1hprtv3do.md +++ b/src/data/roadmaps/python-data-analysis/content/lists@F4dp3Ip8G-EI1hprtv3do.md @@ -1,3 +1,8 @@ # Lists - -Lists are ordered, mutable sequences that can hold elements of any type. They are one of the most used data structures in Python for storing collections of values. Typical uses include holding column names, storing results from loops, and passing multiple values to functions. \ No newline at end of file + +Lists are ordered, mutable sequences that can hold elements of any type. They are one of the most used data structures in Python for storing collections of values. Typical uses include holding column names, storing results from loops, and passing multiple values to functions. + +Visit the following resources to learn more: + +- [@official@List](https://docs.python.org/3/tutorial/datastructures.html) +- [@video@How to Use Lists in Python](https://www.youtube.com/watch?v=9OeznAkyQz4) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/logical@_0K6o-R0F2a9bpEDfRcTi.md b/src/data/roadmaps/python-data-analysis/content/logical@_0K6o-R0F2a9bpEDfRcTi.md index 95d6b8d9a..bfd2012ec 100644 --- a/src/data/roadmaps/python-data-analysis/content/logical@_0K6o-R0F2a9bpEDfRcTi.md +++ b/src/data/roadmaps/python-data-analysis/content/logical@_0K6o-R0F2a9bpEDfRcTi.md @@ -1,3 +1,9 @@ # Logical - -Logical operators combine boolean expressions. Python uses `and`, `or`, and `not` for this purpose. They are used heavily in data filtering conditions, such as selecting rows where multiple criteria are true simultaneously. \ No newline at end of file + +Logical operators combine boolean expressions. Python uses `and`, `or`, and `not` for this purpose. They are used heavily in data filtering conditions, such as selecting rows where multiple criteria are true simultaneously. + +Visit the following resources to learn more: + +- [@article@Python Logical Operators](https://www.w3schools.com/python/python_if_logical.asp) +- [@article@Python not Operator: The Complete Guide to Logical Negation](https://roadmap.sh/python/not-operator) +- [@video@Logical operators in Python are easy πŸ”£](https://www.youtube.com/watch?v=W7luvtXeQTA) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/loops@Dvy7BnNzK55qbh_SgOk8m.md b/src/data/roadmaps/python-data-analysis/content/loops@Dvy7BnNzK55qbh_SgOk8m.md index fe0f76a1f..2806c9e39 100644 --- a/src/data/roadmaps/python-data-analysis/content/loops@Dvy7BnNzK55qbh_SgOk8m.md +++ b/src/data/roadmaps/python-data-analysis/content/loops@Dvy7BnNzK55qbh_SgOk8m.md @@ -1,3 +1,10 @@ # Loops - -Loops execute a block of code repeatedly. Python provides `for` loops for iterating over sequences and `while` loops for condition-based repetition. They are used for batch processing files, iterating over grouped data, and automating repetitive tasks, though vectorized operations are preferred for performance. \ No newline at end of file + +Loops execute a block of code repeatedly. Python provides `for` loops for iterating over sequences and `while` loops for condition-based repetition. They are used for batch processing files, iterating over grouped data, and automating repetitive tasks, though vectorized operations are preferred for performance. + +Visit the following resources to learn more: + +- [@article@Python while Loops: Repeating Tasks Conditionally](https://realpython.com/python-while-loop/) +- [@article@Python for Loops: The Pythonic Way](https://realpython.com/python-for-loop/#the-guts-of-the-python-for-loop) +- [@video@Learn Python for loops in 5 minutes!](https://www.youtube.com/watch?v=KWgYha0clzw) +- [@video@While loops in Python are easy! ♾️](https://www.youtube.com/watch?v=rRTjPnVooxE) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/matplotlib@Md9Yq7bwDNo6hO_A9aa9N.md b/src/data/roadmaps/python-data-analysis/content/matplotlib@Md9Yq7bwDNo6hO_A9aa9N.md index 5b997d021..df5368a04 100644 --- a/src/data/roadmaps/python-data-analysis/content/matplotlib@Md9Yq7bwDNo6hO_A9aa9N.md +++ b/src/data/roadmaps/python-data-analysis/content/matplotlib@Md9Yq7bwDNo6hO_A9aa9N.md @@ -1,3 +1,9 @@ # Matplotlib - -Matplotlib is Python's foundational plotting library. It provides a MATLAB-like interface for creating static, animated, and interactive visualizations. While more verbose than higher-level libraries, Matplotlib offers the most control over every aspect of a plot and is the basis for understanding how other Python visualization tools work. \ No newline at end of file + +Matplotlib is Python's foundational plotting library. It provides a MATLAB-like interface for creating static, animated, and interactive visualizations. While more verbose than higher-level libraries, Matplotlib offers the most control over every aspect of a plot and is the basis for understanding how other Python visualization tools work. + +Visit the following resources to learn more: + +- [@official@Matplotlib Tutorials](https://matplotlib.org/stable/tutorials/index.html) +- [@article@Matplotlib Tutorial](https://www.w3schools.com/python/matplotlib_intro.asp) +- [@video@Matplotlib Full Python Course - Data Science Fundamentals](https://www.youtube.com/watch?v=OZOOLe2imFo) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/mean-median-mode@DdA2gJKpgOHj9mqZL-c4F.md b/src/data/roadmaps/python-data-analysis/content/mean-median-mode@DdA2gJKpgOHj9mqZL-c4F.md index 5a4ae1ac4..d029eebd0 100644 --- a/src/data/roadmaps/python-data-analysis/content/mean-median-mode@DdA2gJKpgOHj9mqZL-c4F.md +++ b/src/data/roadmaps/python-data-analysis/content/mean-median-mode@DdA2gJKpgOHj9mqZL-c4F.md @@ -1,3 +1,8 @@ # Mean, Median, Mode - -Mean, median, and mode are measures of central tendency that describe the typical value in a distribution. Pandas computes these with `mean()`, `median()`, and `mode()` on Series or DataFrame columns. Comparing them reveals distribution shape: in a symmetric distribution they are equal; in a skewed one they diverge. \ No newline at end of file + +Mean, median, and mode are measures of central tendency that describe the typical value in a distribution. Pandas computes these with `mean()`, `median()`, and `mode()` on Series or DataFrame columns. Comparing them reveals distribution shape: in a symmetric distribution they are equal; in a skewed one they diverge. + +Visit the following resources to learn more: + +- [@article@A Guide to Metrics in Exploratory Data Analysis](https://towardsdatascience.com/a-guide-to-metrics-in-exploratory-data-analysis-250b33f72297/) +- [@article@Understanding Your Data: The Essentials of Exploratory Data Analysis](https://dev.to/nderitugichuki/understanding-your-data-the-essentials-of-exploratory-data-analysis-400i) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/merging--joining@Rq9TFEEJnjEgVa1DiJxFz.md b/src/data/roadmaps/python-data-analysis/content/merging--joining@Rq9TFEEJnjEgVa1DiJxFz.md index 19e657e82..4539805bd 100644 --- a/src/data/roadmaps/python-data-analysis/content/merging--joining@Rq9TFEEJnjEgVa1DiJxFz.md +++ b/src/data/roadmaps/python-data-analysis/content/merging--joining@Rq9TFEEJnjEgVa1DiJxFz.md @@ -1,3 +1,8 @@ # Merging & Joining - -Pandas provides `merge()` and `join()` for combining DataFrames based on common columns or indices. Merge supports inner, left, right, and outer joins, mirroring SQL JOIN behavior. `concat()` stacks DataFrames vertically or horizontally. These operations are used to combine data from multiple sources into a single analysis-ready table. \ No newline at end of file + +Pandas provides `merge()` and `join()` for combining DataFrames based on common columns or indices. Merge supports inner, left, right, and outer joins, mirroring SQL JOIN behavior. `concat()` stacks DataFrames vertically or horizontally. These operations are used to combine data from multiple sources into a single analysis-ready table. + +Visit the following resources to learn more: + +- [@official@Merge, join, concatenate and compare](https://pandas.pydata.org/docs/user_guide/merging.html) +- [@article@Combining Data in pandas With merge(), .join(), and concat()](https://realpython.com/pandas-merge-join-and-concat/) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/numpy@IO-disDSfkPau4Pd_TNwc.md b/src/data/roadmaps/python-data-analysis/content/numpy@IO-disDSfkPau4Pd_TNwc.md index e81e85e01..54c5df5e0 100644 --- a/src/data/roadmaps/python-data-analysis/content/numpy@IO-disDSfkPau4Pd_TNwc.md +++ b/src/data/roadmaps/python-data-analysis/content/numpy@IO-disDSfkPau4Pd_TNwc.md @@ -1,3 +1,9 @@ # NumPy - -NumPy is the foundational numerical computing library for Python. It provides the `ndarray`, a fast, multi-dimensional array, and a comprehensive library of mathematical functions that operate on arrays without Python loops. NumPy underpins Pandas, Scikit-learn, and most other scientific Python libraries. \ No newline at end of file + +NumPy is the foundational numerical computing library for Python. It provides the `ndarray`, a fast, multi-dimensional array, and a comprehensive library of mathematical functions that operate on arrays without Python loops. NumPy underpins Pandas, Scikit-learn, and most other scientific Python libraries. + +Visit the following resources to learn more: + +- [@official@NumPy quickstart](https://numpy.org/doc/stable/user/quickstart.html) +- [@article@NumPy Tutorial](https://www.w3schools.com/python/numpy/default.asp) +- [@video@Python NumPy Tutorial for Beginners](https://www.youtube.com/watch?v=QUT1VHiLmmI) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/oop-for-data-analysis@hTAQSe0rHv7fuAMfxJwGp.md b/src/data/roadmaps/python-data-analysis/content/oop-for-data-analysis@hTAQSe0rHv7fuAMfxJwGp.md index 617ea7f42..854217feb 100644 --- a/src/data/roadmaps/python-data-analysis/content/oop-for-data-analysis@hTAQSe0rHv7fuAMfxJwGp.md +++ b/src/data/roadmaps/python-data-analysis/content/oop-for-data-analysis@hTAQSe0rHv7fuAMfxJwGp.md @@ -1,3 +1,9 @@ # OOP for Data Analysis -Object-oriented programming (OOP) organizes code around classes and objects rather than standalone functions and procedures. A class defines a blueprint with attributes (data) and methods (behavior), and objects are instances of that class. For data analysis work, OOP is useful when building reusable data processing components, custom dataset loaders, or analysis pipelines that need to maintain state across multiple steps. Most of the libraries used daily, including Pandas, NumPy, and Scikit-learn, are built around classes, so understanding OOP helps in reading documentation, subclassing existing components, and writing cleaner, more maintainable analysis code. \ No newline at end of file +Object-oriented programming (OOP) organizes code around classes and objects rather than standalone functions and procedures. A class defines a blueprint with attributes (data) and methods (behavior), and objects are instances of that class. For data analysis work, OOP is useful when building reusable data processing components, custom dataset loaders, or analysis pipelines that need to maintain state across multiple steps. Most of the libraries used daily, including Pandas, NumPy, and Scikit-learn, are built around classes, so understanding OOP helps in reading documentation, subclassing existing components, and writing cleaner, more maintainable analysis code. + +Visit the following resources to learn more: + +- [@article@Object-Oriented Programming (OOP) in Python](https://towardsdatascience.com/object-oriented-programming-oop-in-python-56b1f3229c0f/) +- [@article@How data scientists can leverage object oriented programming (OOP)](https://medium.com/@lawjimmy123/how-data-scientists-can-leverage-object-oriented-programming-oop-design-pattern-to-write-better-699166910882) +- [@article@Object-Oriented Programming (OOP) in Python](https://realpython.com/python3-object-oriented-programming/) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/operators@so95CO6Qw3I0S98ISENS-.md b/src/data/roadmaps/python-data-analysis/content/operators@so95CO6Qw3I0S98ISENS-.md index 29a1cfc12..551406e36 100644 --- a/src/data/roadmaps/python-data-analysis/content/operators@so95CO6Qw3I0S98ISENS-.md +++ b/src/data/roadmaps/python-data-analysis/content/operators@so95CO6Qw3I0S98ISENS-.md @@ -1,3 +1,9 @@ # Operators - -Operators are symbols that perform operations on values and variables. Python supports arithmetic, comparison, and logical operators, each serving a different purpose in data analysis code. Understanding how operators work and combine is necessary for writing correct filtering conditions, calculations, and control flow logic. \ No newline at end of file + +Operators are symbols that perform operations on values and variables. Python supports arithmetic, comparison, and logical operators, each serving a different purpose in data analysis code. Understanding how operators work and combine is necessary for writing correct filtering conditions, calculations, and control flow logic. + +Visit the following resources to learn more: + +- [@article@Python Operators](https://www.w3schools.com/python/python_operators.asp) +- [@article@Python Operators with examples](https://www.programiz.com/python-programming/operators) +- [@video@Python Tutorial for Beginners | Operators in Python](https://www.youtube.com/watch?v=v5MR5JnKcZI) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/pandas-string-methods@yRHodYYqdICLZyG3J9qjc.md b/src/data/roadmaps/python-data-analysis/content/pandas-string-methods@yRHodYYqdICLZyG3J9qjc.md index 8b554a147..4b5fa96ca 100644 --- a/src/data/roadmaps/python-data-analysis/content/pandas-string-methods@yRHodYYqdICLZyG3J9qjc.md +++ b/src/data/roadmaps/python-data-analysis/content/pandas-string-methods@yRHodYYqdICLZyG3J9qjc.md @@ -1,3 +1,9 @@ # Pandas String Methods - -Pandas exposes string methods through the `.str` accessor on Series, allowing vectorized text operations on entire columns. Methods include `.str.strip()`, `.str.lower()`, `.str.contains()`, `.str.replace()`, `.str.split()`, and `.str.extract()`. These methods avoid the need to loop over rows for string cleaning. \ No newline at end of file + +Pandas exposes string methods through the `.str` accessor on Series, allowing vectorized text operations on entire columns. Methods include `.str.strip()`, `.str.lower()`, `.str.contains()`, `.str.replace()`, `.str.split()`, and `.str.extract()`. These methods avoid the need to loop over rows for string cleaning. + +Visit the following resources to learn more: + +- [@official@Working with text data](https://pandas.pydata.org/docs/user_guide/text.html) +- [@article@5 Must-Know Pandas Operations on Strings](https://towardsdatascience.com/5-must-know-pandas-operations-on-strings-4f88ca6b8e25/) +- [@video@How do I use string methods in pandas?](https://www.youtube.com/watch?v=bofaC0IckHo) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/pandas@bIblb5cVNSZLqSuskmKnP.md b/src/data/roadmaps/python-data-analysis/content/pandas@bIblb5cVNSZLqSuskmKnP.md index be0711f54..885c35068 100644 --- a/src/data/roadmaps/python-data-analysis/content/pandas@bIblb5cVNSZLqSuskmKnP.md +++ b/src/data/roadmaps/python-data-analysis/content/pandas@bIblb5cVNSZLqSuskmKnP.md @@ -1,3 +1,9 @@ # Pandas - -Pandas is the primary data manipulation library for Python. It provides two core data structures: Series (one-dimensional) and DataFrame (two-dimensional tabular data). Pandas supports loading data from many formats, cleaning, transforming, grouping, merging, and exporting data, covering the full data preparation workflow. \ No newline at end of file + +Pandas is the primary data manipulation library for Python. It provides two core data structures: Series (one-dimensional) and DataFrame (two-dimensional tabular data). Pandas supports loading data from many formats, cleaning, transforming, grouping, merging, and exporting data, covering the full data preparation workflow. + +Visit the following resources to learn more: + +- [@official@10 minutes to pandas](https://pandas.pydata.org/docs/user_guide/10min.html) +- [@article@Pandas Tutorial](https://www.w3schools.com/python/pandas/) +- [@video@Learn Pandas in 1 hour! 🐼](https://www.youtube.com/watch?v=VXtjG_GzO7Q) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/pandas@cD139cRk1xj2ip-o_078l.md b/src/data/roadmaps/python-data-analysis/content/pandas@cD139cRk1xj2ip-o_078l.md index 1ea28dc42..55b492821 100644 --- a/src/data/roadmaps/python-data-analysis/content/pandas@cD139cRk1xj2ip-o_078l.md +++ b/src/data/roadmaps/python-data-analysis/content/pandas@cD139cRk1xj2ip-o_078l.md @@ -1,3 +1,8 @@ # Pandas - -Pandas integrates with SQL through `pd.read_sql()`, which executes a SQL query against a database connection and returns the result as a DataFrame. This allows analysts to leverage SQL for initial data extraction and filtering while using Pandas for downstream manipulation and analysis. \ No newline at end of file + +Pandas integrates with SQL through `pd.read_sql()`, which executes a SQL query against a database connection and returns the result as a DataFrame. This allows analysts to leverage SQL for initial data extraction and filtering while using Pandas for downstream manipulation and analysis. + +Visit the following resources to learn more: + +- [@official@pandas.read_sql](https://pandas.pydata.org/docs/reference/api/pandas.read_sql.html) +- [@video@Read Write Data From Database (read_sql, to_sql)](https://www.youtube.com/watch?v=M-4EpNdlSuY) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/parquet@sBKhhjfFTpbj7TQwVnX7F.md b/src/data/roadmaps/python-data-analysis/content/parquet@sBKhhjfFTpbj7TQwVnX7F.md index e11f7de77..a3442082e 100644 --- a/src/data/roadmaps/python-data-analysis/content/parquet@sBKhhjfFTpbj7TQwVnX7F.md +++ b/src/data/roadmaps/python-data-analysis/content/parquet@sBKhhjfFTpbj7TQwVnX7F.md @@ -1,3 +1,8 @@ # Parquet - -Parquet is a columnar file format optimized for analytical workloads. It stores data with type information, supports efficient compression, and reads much faster than CSV for large datasets. `pd.read_parquet()` requires the `pyarrow` or `fastparquet` library and is the preferred format for storing processed DataFrames on disk. \ No newline at end of file + +Parquet is a columnar file format optimized for analytical workloads. It stores data with type information, supports efficient compression, and reads much faster than CSV for large datasets. `pd.read_parquet()` requires the `pyarrow` or `fastparquet` library and is the preferred format for storing processed DataFrames on disk. + +Visit the following resources to learn more: + +- [@official@pandas.read_parquet](https://pandas.pydata.org/docs/reference/api/pandas.read_parquet.html) +- [@video@Reading Parquet Files in Python](https://www.youtube.com/watch?v=XFO5jdGsMek) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/parsing-dates@L7btVRwvkEVX0K8LxnbqJ.md b/src/data/roadmaps/python-data-analysis/content/parsing-dates@L7btVRwvkEVX0K8LxnbqJ.md index 97063af6a..c46c255a7 100644 --- a/src/data/roadmaps/python-data-analysis/content/parsing-dates@L7btVRwvkEVX0K8LxnbqJ.md +++ b/src/data/roadmaps/python-data-analysis/content/parsing-dates@L7btVRwvkEVX0K8LxnbqJ.md @@ -1,3 +1,9 @@ # Parsing Dates - -Date columns loaded from CSV are typically read as strings and must be converted to datetime objects for time-based operations. `pd.to_datetime()` parses date strings in many formats and accepts a `format` parameter for custom patterns. Once parsed, datetime columns enable operations like extracting year/month, computing differences, and resampling time series. \ No newline at end of file + +Date columns loaded from CSV are typically read as strings and must be converted to datetime objects for time-based operations. `pd.to_datetime()` parses date strings in many formats and accepts a `format` parameter for custom patterns. Once parsed, datetime columns enable operations like extracting year/month, computing differences, and resampling time series. + +Visit the following resources to learn more: + +- [@official@pandas.to_datetime](https://pandas.pydata.org/docs/reference/api/pandas.to_datetime.html) +- [@article@Working with Dates and Time Series Data](https://www.youtube.com/watch?v=UFuo7EHI8zc) +- [@article@Dealing with Date and Time in Pandas DataFrames](https://towardsdatascience.com/dealing-with-date-and-time-in-pandas-dataframes-7d140f711a47/) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/pip@6dJaQRj2zMz00sfTiTdww.md b/src/data/roadmaps/python-data-analysis/content/pip@6dJaQRj2zMz00sfTiTdww.md index c5b5d01c5..11c17a567 100644 --- a/src/data/roadmaps/python-data-analysis/content/pip@6dJaQRj2zMz00sfTiTdww.md +++ b/src/data/roadmaps/python-data-analysis/content/pip@6dJaQRj2zMz00sfTiTdww.md @@ -1,3 +1,9 @@ # pip - -pip is Python's default package installer. It installs packages from the Python Package Index (PyPI) using `pip install package-name`. Libraries like NumPy, Pandas, Matplotlib, and Scikit-learn are all installed this way. A `requirements.txt` file captures all dependencies for a project. \ No newline at end of file + +pip is Python's default package installer. It installs packages from the Python Package Index (PyPI) using `pip install package-name`. Libraries like NumPy, Pandas, Matplotlib, and Scikit-learn are all installed this way. A `requirements.txt` file captures all dependencies for a project. + +Visit the following resources to learn more: + +- [@official@Installing Packages](https://packaging.python.org/en/latest/tutorials/installing-packages/) +- [@opensource@pip](https://github.com/pypa/pip) +- [@video@Python pip πŸ—οΈ](https://www.youtube.com/watch?v=9z7gGUbAj5U) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/plot-categories@5HEYmbVJ6Z7lvgqalVjxW.md b/src/data/roadmaps/python-data-analysis/content/plot-categories@5HEYmbVJ6Z7lvgqalVjxW.md index 529f2004f..b861cd2ef 100644 --- a/src/data/roadmaps/python-data-analysis/content/plot-categories@5HEYmbVJ6Z7lvgqalVjxW.md +++ b/src/data/roadmaps/python-data-analysis/content/plot-categories@5HEYmbVJ6Z7lvgqalVjxW.md @@ -1,3 +1,9 @@ # Plot Categories - -Matplotlib supports a wide range of plot types: line plots (`plot`), bar charts (`bar`, `barh`), scatter plots (`scatter`), histograms (`hist`), box plots (`boxplot`), pie charts (`pie`), and more. Choosing the right plot type depends on the data structure and the relationship being communicated. \ No newline at end of file + +Matplotlib supports a wide range of plot types: line plots (`plot`), bar charts (`bar`, `barh`), scatter plots (`scatter`), histograms (`hist`), box plots (`boxplot`), pie charts (`pie`), and more. Choosing the right plot type depends on the data structure and the relationship being communicated. + +Visit the following resources to learn more: + +- [@official@Plot types](https://matplotlib.org/stable/plot_types/index.html) +- [@article@Python Gallery: Matplotlib](https://python-graph-gallery.com/matplotlib/) +- [@article@Matplotlib: Part 3. Exploring Different Plot Types](https://medium.com/@ebimsv/mastering-matplotlib-3-exploring-different-plot-types-bd13d18ff613) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/plotly@J_u8yvrHEWKKmVsVgPN7L.md b/src/data/roadmaps/python-data-analysis/content/plotly@J_u8yvrHEWKKmVsVgPN7L.md index 83f421cb1..39e7dfa35 100644 --- a/src/data/roadmaps/python-data-analysis/content/plotly@J_u8yvrHEWKKmVsVgPN7L.md +++ b/src/data/roadmaps/python-data-analysis/content/plotly@J_u8yvrHEWKKmVsVgPN7L.md @@ -1,3 +1,9 @@ # Plotly - -Plotly is a Python library for creating interactive charts and dashboards. It produces web-based visualizations using JavaScript under the hood, with a Python API. Plotly Express provides a high-level interface for common chart types, while the `graph_objects` module offers full control. Plotly integrates with Dash for building full web dashboards. \ No newline at end of file + +Plotly is a Python library for creating interactive charts and dashboards. It produces web-based visualizations using JavaScript under the hood, with a Python API. Plotly Express provides a high-level interface for common chart types, while the `graph_objects` module offers full control. Plotly integrates with Dash for building full web dashboards. + +Visit the following resources to learn more: + +- [@official@Getting Started with Plotly in Python](https://plotly.com/python/getting-started/) +- [@official@Plotly Python Graphing Library Fundamentals](https://plotly.com/python/plotly-fundamentals/) +- [@video@Plotly Tutorial - Basics in 7 Minutes!](https://www.youtube.com/watch?v=PqUaDvbczbI) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/polars@18eitPvBQKdeM3K33d59g.md b/src/data/roadmaps/python-data-analysis/content/polars@18eitPvBQKdeM3K33d59g.md index 9758845be..64a9c8e80 100644 --- a/src/data/roadmaps/python-data-analysis/content/polars@18eitPvBQKdeM3K33d59g.md +++ b/src/data/roadmaps/python-data-analysis/content/polars@18eitPvBQKdeM3K33d59g.md @@ -1,3 +1,8 @@ # Polars - -Polars is a fast DataFrame library for Python written in Rust. It is designed as a high-performance alternative to Pandas, with a more consistent API and significantly better performance on large datasets. Polars uses lazy evaluation and query optimization to process data efficiently without loading everything into memory at once. \ No newline at end of file + +Polars is a fast DataFrame library for Python written in Rust. It is designed as a high-performance alternative to Pandas, with a more consistent API and significantly better performance on large datasets. Polars uses lazy evaluation and query optimization to process data efficiently without loading everything into memory at once. + +Visit the following resources to learn more: + +- [@official@Getting started](https://docs.pola.rs/user-guide/getting-started/) +- [@video@Learning the Polars DataFrame Library!](https://www.youtube.com/watch?v=OTVDmA6CRlQ) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/power-bi--tableau@EHVXO79NGpLC1i8x4Y7g1.md b/src/data/roadmaps/python-data-analysis/content/power-bi--tableau@EHVXO79NGpLC1i8x4Y7g1.md index 654a0177a..88e262be7 100644 --- a/src/data/roadmaps/python-data-analysis/content/power-bi--tableau@EHVXO79NGpLC1i8x4Y7g1.md +++ b/src/data/roadmaps/python-data-analysis/content/power-bi--tableau@EHVXO79NGpLC1i8x4Y7g1.md @@ -1,3 +1,8 @@ # Power BI / Tableau - -Power BI and Tableau are enterprise BI platforms for building interactive dashboards and reports. Python integrates with both: Power BI supports Python visuals and data transformation scripts, and Tableau supports Python through TabPy for custom calculations. Analysts who prepare data in Python can visualize and distribute it through these platforms for business audiences. \ No newline at end of file + +Power BI and Tableau are enterprise BI platforms for building interactive dashboards and reports. Python integrates with both: Power BI supports Python visuals and data transformation scripts, and Tableau supports Python through TabPy for custom calculations. Analysts who prepare data in Python can visualize and distribute it through these platforms for business audiences. + +Visit the following resources to learn more: + +- [@video@Power BI Tutorials for Beginners](https://www.youtube.com/playlist?list=PLUaB-1hjhk8HqnmK0gQhfmIdCbxwoAoys) +- [@video@Learn Tableau in Under 2 hours](https://www.youtube.com/watch?v=j8FSP8XuFyk) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/printing-variables@icVUQgedEGcaPpJSV8JdI.md b/src/data/roadmaps/python-data-analysis/content/printing-variables@icVUQgedEGcaPpJSV8JdI.md index bfce47b86..257fbe2bc 100644 --- a/src/data/roadmaps/python-data-analysis/content/printing-variables@icVUQgedEGcaPpJSV8JdI.md +++ b/src/data/roadmaps/python-data-analysis/content/printing-variables@icVUQgedEGcaPpJSV8JdI.md @@ -1,3 +1,10 @@ # Printing Variables - -Printing variables is done with Python's built-in `print()` function. During analysis, printing intermediate values helps verify that transformations are working as expected. F-strings (`f"value: {variable}"`) provide a clean way to format output with variable values embedded in strings. \ No newline at end of file + +Printing variables is done with Python's built-in `print()` function. During analysis, printing intermediate values helps verify that transformations are working as expected. F-strings (`f"value: {variable}"`) provide a clean way to format output with variable values embedded in strings. + +Visit the following resources to learn more: + +- [@article@Variables in Python: Usage and Best Practices](https://realpython.com/python-variables/) +- [@article@Python Print New Line: Methods, Examples, and Best Practices](https://roadmap.sh/python/print-new-line) +- [@article@Python Multiline Strings: The Complete Guide](https://roadmap.sh/python/multiline-strings) +- [@video@Data Types & Variables in Python](https://www.youtube.com/playlist?list=PLBlnK6fEyqRhN-sfWgCU1z_Qhakc1AGOn) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/pyspark@XTXm8aIDbRkTXoAck_m3I.md b/src/data/roadmaps/python-data-analysis/content/pyspark@XTXm8aIDbRkTXoAck_m3I.md index 121950808..2f737e27b 100644 --- a/src/data/roadmaps/python-data-analysis/content/pyspark@XTXm8aIDbRkTXoAck_m3I.md +++ b/src/data/roadmaps/python-data-analysis/content/pyspark@XTXm8aIDbRkTXoAck_m3I.md @@ -1,3 +1,9 @@ # PySpark - -PySpark is the Python API for Apache Spark, the distributed data processing engine. It allows Python code to run Spark jobs on clusters, processing datasets at the scale of terabytes. PySpark provides DataFrame and SQL APIs similar to Pandas and integrates with MLlib for distributed machine learning. It is used when data volume exceeds what Dask or a single machine can handle. \ No newline at end of file + +PySpark is the Python API for Apache Spark, the distributed data processing engine. It allows Python code to run Spark jobs on clusters, processing datasets at the scale of terabytes. PySpark provides DataFrame and SQL APIs similar to Pandas and integrates with MLlib for distributed machine learning. It is used when data volume exceeds what Dask or a single machine can handle. + +Visit the following resources to learn more: + +- [@official@PySpark Tutorials](https://spark.apache.org/docs/latest/api/python/tutorial/index.html) +- [@article@PySpark for Beginners: Beyond the Basics](https://towardsdatascience.com/pyspark-for-beginners-beyond-the-basics/) +- [@video@PySpark Tutorial](https://www.youtube.com/watch?v=wNRjR6Cds5s&list=PL2IsFZBGM_IHCl9zhRVC1EXTomkEp_1zm) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/random-module@XizmE-QroU4GJbjvqNYxe.md b/src/data/roadmaps/python-data-analysis/content/random-module@XizmE-QroU4GJbjvqNYxe.md index b70920a27..32f1f4a70 100644 --- a/src/data/roadmaps/python-data-analysis/content/random-module@XizmE-QroU4GJbjvqNYxe.md +++ b/src/data/roadmaps/python-data-analysis/content/random-module@XizmE-QroU4GJbjvqNYxe.md @@ -1,3 +1,9 @@ # Random Module - -NumPy's `random` module generates pseudorandom numbers and samples. It provides functions for creating random arrays, sampling from distributions (normal, uniform, binomial), and setting a seed for reproducibility. Random number generation is used in simulation, bootstrapping, and initializing machine learning models. \ No newline at end of file + +NumPy's `random` module generates pseudorandom numbers and samples. It provides functions for creating random arrays, sampling from distributions (normal, uniform, binomial), and setting a seed for reproducibility. Random number generation is used in simulation, bootstrapping, and initializing machine learning models. + +Visit the following resources to learn more: + +- [@article@Random Numbers in NumPy](https://www.w3schools.com/python/numpy/numpy_random.asp) +- [@article@Numpy Random](https://www.programiz.com/python-programming/numpy/random) +- [@video@Random numbers in NumPy are easy! 🎲](https://www.youtube.com/watch?v=Ql5zGPtxlHY&pp=ygUNIG51bXB5IHJhbmRvbQ%3D%3D) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/re@oXhGEIie-MzfUSM2Ng4Zm.md b/src/data/roadmaps/python-data-analysis/content/re@oXhGEIie-MzfUSM2Ng4Zm.md index 467c7fe56..ab938ae8b 100644 --- a/src/data/roadmaps/python-data-analysis/content/re@oXhGEIie-MzfUSM2Ng4Zm.md +++ b/src/data/roadmaps/python-data-analysis/content/re@oXhGEIie-MzfUSM2Ng4Zm.md @@ -1,3 +1,10 @@ # re - -The `re` module provides regular expression support for pattern matching and text manipulation. It is used for extracting structured data from unstructured text, validating formats, and performing complex find-and-replace operations. Key functions include `re.match()`, `re.search()`, `re.findall()`, and `re.sub()`. \ No newline at end of file + +The `re` module provides regular expression support for pattern matching and text manipulation. It is used for extracting structured data from unstructured text, validating formats, and performing complex find-and-replace operations. Key functions include `re.match()`, `re.search()`, `re.findall()`, and `re.sub()`. + +Visit the following resources to learn more: + +- [@official@re β€” Regular expression operations](https://docs.python.org/3/library/re.html) +- [@official@Regular expression HOWTO](https://docs.python.org/3/howto/regex.html) +- [@article@Python RegEx](https://www.w3schools.com/python/python_regex.asp) +- [@video@Python Tutorial: re Module - How to Write and Match Regular Expressions (Regex)](https://www.youtube.com/watch?v=K8L6KVGG-7o) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/reading-data@w2j11Jyr_qLP1Y9ST-06O.md b/src/data/roadmaps/python-data-analysis/content/reading-data@w2j11Jyr_qLP1Y9ST-06O.md index 97b211448..115c013b7 100644 --- a/src/data/roadmaps/python-data-analysis/content/reading-data@w2j11Jyr_qLP1Y9ST-06O.md +++ b/src/data/roadmaps/python-data-analysis/content/reading-data@w2j11Jyr_qLP1Y9ST-06O.md @@ -1,3 +1,9 @@ # Reading Data - -Pandas provides functions to load data from many formats: `pd.read_csv()`, `pd.read_excel()`, `pd.read_json()`, `pd.read_parquet()`, `pd.read_sql()`, and others. Each function returns a DataFrame and accepts parameters for handling headers, delimiters, encoding, and data types. Reading data is always the first step in a Pandas workflow. \ No newline at end of file + +Pandas provides functions to load data from many formats: `pd.read_csv()`, `pd.read_excel()`, `pd.read_json()`, `pd.read_parquet()`, `pd.read_sql()`, and others. Each function returns a DataFrame and accepts parameters for handling headers, delimiters, encoding, and data types. Reading data is always the first step in a Pandas workflow. + +Visit the following resources to learn more: + +- [@official@How do I read and write tabular data?](https://pandas.pydata.org/docs/getting_started/intro_tutorials/02_read_write.html) +- [@article@Pandas 101: How to Read Data from Multiple Sources](https://riyoma.medium.com/pandas-101-how-to-read-data-from-multiple-sources-ba75e5497ad5) +- [@video@Reading in Files in Pandas | Python Pandas Tutorials](https://www.youtube.com/watch?v=dUpyC40cF6Q) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/reading-local-files@34DL5RWQtHkvgVhI9a_ZQ.md b/src/data/roadmaps/python-data-analysis/content/reading-local-files@34DL5RWQtHkvgVhI9a_ZQ.md index 490043bde..14ca80fb1 100644 --- a/src/data/roadmaps/python-data-analysis/content/reading-local-files@34DL5RWQtHkvgVhI9a_ZQ.md +++ b/src/data/roadmaps/python-data-analysis/content/reading-local-files@34DL5RWQtHkvgVhI9a_ZQ.md @@ -1,3 +1,8 @@ # Reading Local Files - -Reading local files loads data stored on disk into Python for analysis. Pandas supports the most common file formats used in data work. The right function to use depends on the file format, and parameters like delimiter, encoding, and header row often need to be specified. \ No newline at end of file + +Reading local files loads data stored on disk into Python for analysis. Pandas supports the most common file formats used in data work. The right function to use depends on the file format, and parameters like delimiter, encoding, and header row often need to be specified. + +Visit the following resources to learn more: + +- [@article@pandas: How to Read and Write Files](https://realpython.com/pandas-read-write-files/) +- [@video@How to Navigate File Paths: Reading Data With Pandas](https://www.youtube.com/watch?v=39pwKSJ7T1Y) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/reading-web-data@w8iQWthKu1w75z36B8Ir3.md b/src/data/roadmaps/python-data-analysis/content/reading-web-data@w8iQWthKu1w75z36B8Ir3.md index ba7989897..1049e0994 100644 --- a/src/data/roadmaps/python-data-analysis/content/reading-web-data@w8iQWthKu1w75z36B8Ir3.md +++ b/src/data/roadmaps/python-data-analysis/content/reading-web-data@w8iQWthKu1w75z36B8Ir3.md @@ -1,3 +1,7 @@ # Reading Web Data - -Reading web data involves fetching data from URLs, REST APIs, and web pages directly into Python. This allows analysis workflows to incorporate live or frequently updated data without manual downloads. The main tools are the `requests` library for APIs and `BeautifulSoup` or `scrapy` for web scraping. \ No newline at end of file + +Reading web data involves fetching data from URLs, REST APIs, and web pages directly into Python. This allows analysis workflows to incorporate live or frequently updated data without manual downloads. The main tools are the `requests` library for APIs and `BeautifulSoup` or `scrapy` for web scraping. + +Visit the following resources to learn more: + +- [@article@An Efficient Way to Read Data from the Web Directly into Python](https://medium.com/data-science/an-efficient-way-to-read-data-from-the-web-directly-into-python-a526a0b4f4cb) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/regression-plots@PcGJMkmZrl4xc2ADT9dF3.md b/src/data/roadmaps/python-data-analysis/content/regression-plots@PcGJMkmZrl4xc2ADT9dF3.md index 5ee315567..bc4f0d65e 100644 --- a/src/data/roadmaps/python-data-analysis/content/regression-plots@PcGJMkmZrl4xc2ADT9dF3.md +++ b/src/data/roadmaps/python-data-analysis/content/regression-plots@PcGJMkmZrl4xc2ADT9dF3.md @@ -1,3 +1,9 @@ # Regression Plots - -Seaborn's regression plots visualize the relationship between two numeric variables with a fitted regression line. `sns.regplot()` plots data points and a linear regression fit with confidence interval. `sns.lmplot()` extends this to support faceting by a categorical variable, enabling comparison across groups. \ No newline at end of file + +Seaborn's regression plots visualize the relationship between two numeric variables with a fitted regression line. `sns.regplot()` plots data points and a linear regression fit with confidence interval. `sns.lmplot()` extends this to support faceting by a categorical variable, enabling comparison across groups. + +Visit the following resources to learn more: + +- [@official@Visualizing statistical relationships](https://seaborn.pydata.org/tutorial/relational.html) +- [@article@Seaborn Relplot in Python: Visualising Relationships in Data](https://towardsdatascience.com/seaborn-relplot-in-python-visualising-relationships-in-data-ee39138d53aa/) +- [@video@Seaborn scatter plot | How to make and style a scatterplot in Python seaborn](https://www.youtube.com/watch?v=4yz4cMXCkuw) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/reshaping@ACr__YweG2xOAf56FK8tE.md b/src/data/roadmaps/python-data-analysis/content/reshaping@ACr__YweG2xOAf56FK8tE.md index 6882aaf85..6f46a9854 100644 --- a/src/data/roadmaps/python-data-analysis/content/reshaping@ACr__YweG2xOAf56FK8tE.md +++ b/src/data/roadmaps/python-data-analysis/content/reshaping@ACr__YweG2xOAf56FK8tE.md @@ -1,3 +1,9 @@ # Reshaping - -Reshaping transforms the structure of a DataFrame without changing its data. `pivot()` and `pivot_table()` convert long-format data to wide format. `melt()` does the reverse, converting wide to long. `stack()` and `unstack()` move index levels to columns or vice versa. Reshaping is often needed to prepare data for specific visualizations or models. \ No newline at end of file + +Reshaping transforms the structure of a DataFrame without changing its data. `pivot()` and `pivot_table()` convert long-format data to wide format. `melt()` does the reverse, converting wide to long. `stack()` and `unstack()` move index levels to columns or vice versa. Reshaping is often needed to prepare data for specific visualizations or models. + +Visit the following resources to learn more: + +- [@official@Reshaping and pivot tables](https://pandas.pydata.org/docs/user_guide/reshaping.html) +- [@article@Reshaping a Pandas Dataframe: Long-to-Wide and Vice Versa](https://towardsdatascience.com/reshaping-a-pandas-dataframe-long-to-wide-and-vice-versa-517c7f0995ad/) +- [@video@Pandas - Reshape Dataframe](https://www.youtube.com/watch?v=oY62o-tBHF4&list=PL6icxsSf2sRe2qqFmtZwU_VKChEChzhEP) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/saving-figures@mElhFGGT2GnLg2AkJFETN.md b/src/data/roadmaps/python-data-analysis/content/saving-figures@mElhFGGT2GnLg2AkJFETN.md index 401173446..b68f611a6 100644 --- a/src/data/roadmaps/python-data-analysis/content/saving-figures@mElhFGGT2GnLg2AkJFETN.md +++ b/src/data/roadmaps/python-data-analysis/content/saving-figures@mElhFGGT2GnLg2AkJFETN.md @@ -1,3 +1,8 @@ # Saving figures - -Figures are saved to disk with `plt.savefig('filename.png', dpi=300, bbox_inches='tight')`. Supported formats include PNG, PDF, SVG, and JPEG. Saving high-resolution figures is important when embedding charts in reports or publications. The `bbox_inches='tight'` parameter prevents axis labels from being cut off. \ No newline at end of file + +Figures are saved to disk with `plt.savefig('filename.png', dpi=300, bbox_inches='tight')`. Supported formats include PNG, PDF, SVG, and JPEG. Saving high-resolution figures is important when embedding charts in reports or publications. The `bbox_inches='tight'` parameter prevents axis labels from being cut off. + +Visit the following resources to learn more: + +- [@article@matplotlib.pyplot.savefig](https://matplotlib.org/stable/api/_as_gen/matplotlib.pyplot.savefig.html) +- [@video@How to save a matplotlib figure and fix text cutting off || Matplotlib Tips](https://www.youtube.com/watch?v=C8MT-A7Mvk4) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/scatterplot@4ce2yYLRCSb3azsOMpE-U.md b/src/data/roadmaps/python-data-analysis/content/scatterplot@4ce2yYLRCSb3azsOMpE-U.md index 6cdca6d71..11d4b2d0e 100644 --- a/src/data/roadmaps/python-data-analysis/content/scatterplot@4ce2yYLRCSb3azsOMpE-U.md +++ b/src/data/roadmaps/python-data-analysis/content/scatterplot@4ce2yYLRCSb3azsOMpE-U.md @@ -1,3 +1,9 @@ # Scatterplot - -A scatter plot displays two numeric variables as points on an x-y axis to reveal their relationship. It is used during EDA to detect correlations, clusters, and outliers. A trend line or regression line can be added to show the direction and strength of the linear relationship between the variables. \ No newline at end of file + +A scatter plot displays two numeric variables as points on an x-y axis to reveal their relationship. It is used during EDA to detect correlations, clusters, and outliers. A trend line or regression line can be added to show the direction and strength of the linear relationship between the variables. + +Visit the following resources to learn more: + +- [@article@matplotlib.pyplot.scatter](https://matplotlib.org/stable/api/_as_gen/matplotlib.pyplot.scatter.html) +- [@article@How to Make a Scatter Plot in Python With plt.scatter()](https://realpython.com/visualizing-python-plt-scatter/) +- [@video@Seaborn scatter plot | How to make and style a scatterplot in Python seaborn](https://www.youtube.com/watch?v=4yz4cMXCkuw) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/scikit-learn@npndQC8OIfP1wwn_iswPy.md b/src/data/roadmaps/python-data-analysis/content/scikit-learn@npndQC8OIfP1wwn_iswPy.md index 4e0e4cc92..8f874975f 100644 --- a/src/data/roadmaps/python-data-analysis/content/scikit-learn@npndQC8OIfP1wwn_iswPy.md +++ b/src/data/roadmaps/python-data-analysis/content/scikit-learn@npndQC8OIfP1wwn_iswPy.md @@ -1,3 +1,8 @@ # Scikit-learn - -Scikit-learn is the standard machine learning library for Python. It provides a consistent API for classification, regression, clustering, dimensionality reduction, model selection, and preprocessing. Models are trained with `.fit()`, used to predict with `.predict()`, and evaluated with a suite of metrics. Scikit-learn also provides tools for pipelines, cross-validation, and hyperparameter tuning. \ No newline at end of file + +Scikit-learn is the standard machine learning library for Python. It provides a consistent API for classification, regression, clustering, dimensionality reduction, model selection, and preprocessing. Models are trained with `.fit()`, used to predict with `.predict()`, and evaluated with a suite of metrics. Scikit-learn also provides tools for pipelines, cross-validation, and hyperparameter tuning. + +Visit the following resources to learn more: + +- [@article@Scikit-learn](https://scikit-learn.org/stable/) +- [@video@Scikit-learn Crash Course - Machine Learning Library for Python](https://www.youtube.com/watch?v=0B5eIE_1vpU) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/scipy@YxA0ms9vkufKGOguRNK1B.md b/src/data/roadmaps/python-data-analysis/content/scipy@YxA0ms9vkufKGOguRNK1B.md index 0fd674e71..b7e4735e5 100644 --- a/src/data/roadmaps/python-data-analysis/content/scipy@YxA0ms9vkufKGOguRNK1B.md +++ b/src/data/roadmaps/python-data-analysis/content/scipy@YxA0ms9vkufKGOguRNK1B.md @@ -1,3 +1,9 @@ # SciPy - -SciPy is a scientific computing library built on NumPy. It provides modules for statistics (`scipy.stats`), optimization (`scipy.optimize`), linear algebra, signal processing, and numerical integration. `scipy.stats` is used for hypothesis tests (t-tests, chi-square, ANOVA), probability distributions, and descriptive statistics beyond what NumPy provides. \ No newline at end of file + +SciPy is a scientific computing library built on NumPy. It provides modules for statistics (`scipy.stats`), optimization (`scipy.optimize`), linear algebra, signal processing, and numerical integration. `scipy.stats` is used for hypothesis tests (t-tests, chi-square, ANOVA), probability distributions, and descriptive statistics beyond what NumPy provides. + +Visit the following resources to learn more: + +- [@official@SciPy User Guide](https://docs.scipy.org/doc/scipy/tutorial/) +- [@article@SciPy Tutorial](https://www.w3schools.com/python/scipy/index.php) +- [@video@SciPy Tutorial: For Physicists, Engineers, and Mathematicians](https://www.youtube.com/watch?v=jmX4FOUEfgU) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/scrapy@UImXdBHQKCzp9NdjF9xMi.md b/src/data/roadmaps/python-data-analysis/content/scrapy@UImXdBHQKCzp9NdjF9xMi.md index 705f5cb3f..82d412754 100644 --- a/src/data/roadmaps/python-data-analysis/content/scrapy@UImXdBHQKCzp9NdjF9xMi.md +++ b/src/data/roadmaps/python-data-analysis/content/scrapy@UImXdBHQKCzp9NdjF9xMi.md @@ -1,3 +1,9 @@ # scrapy - -Scrapy is a Python framework for large-scale web scraping. Unlike BeautifulSoup, which parses individual pages, Scrapy manages the full crawling workflow: following links, handling pagination, managing request queues, and exporting data. It is used when scraping requires collecting data from many pages across a site. \ No newline at end of file + +Scrapy is a Python framework for large-scale web scraping. Unlike BeautifulSoup, which parses individual pages, Scrapy manages the full crawling workflow: following links, handling pagination, managing request queues, and exporting data. It is used when scraping requires collecting data from many pages across a site. + +Visit the following resources to learn more: + +- [@official@Scrapy Docs](https://docs.scrapy.org/en/latest/) +- [@article@A Minimalist End-to-End Scrapy Tutorial (Part I)](https://medium.com/data-science/a-minimalist-end-to-end-scrapy-tutorial-part-i-11e350bcdec0) +- [@video@Scrapy for Beginners - A Complete How To Example Web Scraping Project](https://www.youtube.com/watch?v=s4jtkzHhLzY) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/seaborn@zx7uCaQiCxvXMhm9FG3XG.md b/src/data/roadmaps/python-data-analysis/content/seaborn@zx7uCaQiCxvXMhm9FG3XG.md index 6c0c5b6e3..16e07dd2c 100644 --- a/src/data/roadmaps/python-data-analysis/content/seaborn@zx7uCaQiCxvXMhm9FG3XG.md +++ b/src/data/roadmaps/python-data-analysis/content/seaborn@zx7uCaQiCxvXMhm9FG3XG.md @@ -1,3 +1,10 @@ # Seaborn - -Seaborn is a Python visualization library built on Matplotlib that provides a higher-level interface for statistical graphics. It handles common plot types with less code and integrates tightly with Pandas DataFrames. Seaborn is particularly strong for visualizing statistical relationships, distributions, and grouped comparisons. \ No newline at end of file + +Seaborn is a Python visualization library built on Matplotlib that provides a higher-level interface for statistical graphics. It handles common plot types with less code and integrates tightly with Pandas DataFrames. Seaborn is particularly strong for visualizing statistical relationships, distributions, and grouped comparisons. + +Visit the following resources to learn more: + +- [@official@Seaborn Tutorials](https://seaborn.pydata.org/tutorial.html) +- [@official@Introduction to Seaborn for dataviz with Python](https://python-graph-gallery.com/seaborn/) +- [@article@Best Seaborn Visualizations for Data Science](https://towardsdatascience.com/best-seaborn-visualizations-for-data-science-3d866f99c3a9/) +- [@video@Introduction to Seaborn](https://www.youtube.com/watch?v=vaf4ir8eT38&list=PLtPIclEQf-3cG31dxSMZ8KTcDG7zYng1j) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/series-and-dataframe@3N5N6qe8JVr0XGOFdn5-8.md b/src/data/roadmaps/python-data-analysis/content/series-and-dataframe@3N5N6qe8JVr0XGOFdn5-8.md index 03153e502..9a4a13ed4 100644 --- a/src/data/roadmaps/python-data-analysis/content/series-and-dataframe@3N5N6qe8JVr0XGOFdn5-8.md +++ b/src/data/roadmaps/python-data-analysis/content/series-and-dataframe@3N5N6qe8JVr0XGOFdn5-8.md @@ -1,3 +1,10 @@ # Series and DataFrame - -A Series is a one-dimensional labeled array, analogous to a single column in a spreadsheet. A DataFrame is a two-dimensional table of Series that share an index, analogous to a spreadsheet or SQL table. These two structures are the foundation of all Pandas operations. \ No newline at end of file + +A Series is a one-dimensional labeled array, analogous to a single column in a spreadsheet. A DataFrame is a two-dimensional table of Series that share an index, analogous to a spreadsheet or SQL table. These two structures are the foundation of all Pandas operations. + +Visit the following resources to learn more: + +- [@official@pandas.Series](https://pandas.pydata.org/docs/reference/api/pandas.Series.html) +- [@official@pandas.DataFrame](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.html) +- [@article@Pandas Series](https://www.w3schools.com/python/pandas/pandas_series.asp) +- [@article@Pandas DataFrames](https://www.w3schools.com/python/pandas/pandas_dataframes.asp) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/sets@GvV_thqIC0yXyJbJt_Zmu.md b/src/data/roadmaps/python-data-analysis/content/sets@GvV_thqIC0yXyJbJt_Zmu.md index d2b846ab4..658022ed1 100644 --- a/src/data/roadmaps/python-data-analysis/content/sets@GvV_thqIC0yXyJbJt_Zmu.md +++ b/src/data/roadmaps/python-data-analysis/content/sets@GvV_thqIC0yXyJbJt_Zmu.md @@ -1,3 +1,8 @@ # Sets - -Sets are unordered collections of unique values. They support mathematical set operations like union, intersection, and difference. They are useful for finding unique values, checking membership, and comparing two groups of items. \ No newline at end of file + +Sets are unordered collections of unique values. They support mathematical set operations like union, intersection, and difference. They are useful for finding unique values, checking membership, and comparing two groups of items. + +Visit the following resources to learn more: + +- [@official@Sets](https://docs.python.org/3/tutorial/datastructures.html#sets) +- [@video@What are Sets in Python? Python Tutorial for Absolute Beginners](https://www.youtube.com/watch?v=t9j8lCUGZXo) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/sql-fundamentals@E_r4BLrSyHkrxpg4HDkUm.md b/src/data/roadmaps/python-data-analysis/content/sql-fundamentals@E_r4BLrSyHkrxpg4HDkUm.md index df46959a5..ac89d3dfe 100644 --- a/src/data/roadmaps/python-data-analysis/content/sql-fundamentals@E_r4BLrSyHkrxpg4HDkUm.md +++ b/src/data/roadmaps/python-data-analysis/content/sql-fundamentals@E_r4BLrSyHkrxpg4HDkUm.md @@ -1,3 +1,9 @@ # SQL Fundamentals - -SQL (Structured Query Language) is the standard language for querying relational databases. Data analysts use SQL to extract, filter, aggregate, and join data from databases before loading it into Python for further analysis. Python provides several libraries for running SQL queries directly from code. \ No newline at end of file + +SQL (Structured Query Language) is the standard language for querying relational databases. Data analysts use SQL to extract, filter, aggregate, and join data from databases before loading it into Python for further analysis. Python provides several libraries for running SQL queries directly from code. + +Visit the following resources to learn more: + +- [@article@SQL Tutorial - Mode](https://www.thoughtspot.com/sql-tutorial) +- [@article@SQL Tutorial](https://www.sqltutorial.org/) +- [@video@SQL Tutorial - Full Database Course for Beginners](https://www.youtube.com/watch?v=HXV3zeQKqGY) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/sqlalchemy@k63rFXaVDUX2ZOVI_O9WG.md b/src/data/roadmaps/python-data-analysis/content/sqlalchemy@k63rFXaVDUX2ZOVI_O9WG.md index a1b268c58..03960832a 100644 --- a/src/data/roadmaps/python-data-analysis/content/sqlalchemy@k63rFXaVDUX2ZOVI_O9WG.md +++ b/src/data/roadmaps/python-data-analysis/content/sqlalchemy@k63rFXaVDUX2ZOVI_O9WG.md @@ -1,3 +1,9 @@ # SQLAlchemy - -SQLAlchemy is a Python SQL toolkit and object-relational mapper (ORM) that provides a unified interface for connecting to many database backends including PostgreSQL, MySQL, SQLite, and others. It is used with Pandas via `pd.read_sql()` to load query results directly into DataFrames. \ No newline at end of file + +SQLAlchemy is a Python SQL toolkit and object-relational mapper (ORM) that provides a unified interface for connecting to many database backends including PostgreSQL, MySQL, SQLite, and others. It is used with Pandas via `pd.read_sql()` to load query results directly into DataFrames. + +Visit the following resources to learn more: + +- [@official@SQLAlchemy](https://www.sqlalchemy.org/) +- [@article@Mastering SQLAlchemy: A Comprehensive Guide for Python Developers](https://medium.com/@ramanbazhanau/mastering-sqlalchemy-a-comprehensive-guide-for-python-developers-ddb3d9f2e829) +- [@video@SQLAlchemy: The BEST SQL Database Library in Python](https://www.youtube.com/watch?v=aAy-B6KPld8) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/sqlite3@-RepeaHp66GGuM66k9JHd.md b/src/data/roadmaps/python-data-analysis/content/sqlite3@-RepeaHp66GGuM66k9JHd.md index f151f9c62..bcd010158 100644 --- a/src/data/roadmaps/python-data-analysis/content/sqlite3@-RepeaHp66GGuM66k9JHd.md +++ b/src/data/roadmaps/python-data-analysis/content/sqlite3@-RepeaHp66GGuM66k9JHd.md @@ -1,3 +1,8 @@ # sqlite3 - -`sqlite3` is Python's built-in library for working with SQLite databases. SQLite is a lightweight, file-based relational database that requires no server setup. It is commonly used for local data storage, prototyping, and working with small to medium datasets entirely within Python. \ No newline at end of file + +`sqlite3` is Python's built-in library for working with SQLite databases. SQLite is a lightweight, file-based relational database that requires no server setup. It is commonly used for local data storage, prototyping, and working with small to medium datasets entirely within Python. + +Visit the following resources to learn more: + +- [@official@sqlite3](https://docs.python.org/3/library/sqlite3.html#sqlite3-tutorial) +- [@video@Sqlite 3 Python Tutorial in 5 minutes](https://www.youtube.com/watch?v=girsuXz0yA8) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/statistics--ml@KLXFKVb-EC5WOigI63y09.md b/src/data/roadmaps/python-data-analysis/content/statistics--ml@KLXFKVb-EC5WOigI63y09.md index 4e7445cad..ecef9f8d0 100644 --- a/src/data/roadmaps/python-data-analysis/content/statistics--ml@KLXFKVb-EC5WOigI63y09.md +++ b/src/data/roadmaps/python-data-analysis/content/statistics--ml@KLXFKVb-EC5WOigI63y09.md @@ -1,3 +1,9 @@ # Statistics & ML - -Python has a rich ecosystem of libraries for statistical analysis and machine learning. SciPy extends NumPy with statistical tests, optimization, and signal processing. Scikit-learn provides a consistent API for building, evaluating, and deploying machine learning models. Together they cover the analytical needs of most data analysis work. \ No newline at end of file + +Python has a rich ecosystem of libraries for statistical analysis and machine learning. SciPy extends NumPy with statistical tests, optimization, and signal processing. Scikit-learn provides a consistent API for building, evaluating, and deploying machine learning models. Together they cover the analytical needs of most data analysis work. + +Visit the following resources to learn more: + +- [@article@Python Statistics Fundamentals: How to Describe Your Data](https://realpython.com/python-statistics/) +- [@video@Mastering Probability and Statistics in Python](https://www.youtube.com/playlist?list=PLVgEzPHodXi1wT9OK8B_W6Hs8Xc-gaG6N) +- [@video@Machine Learning Tutorial Python | Machine Learning For Beginners](https://www.youtube.com/playlist?list=PLeo1K3hjS3uvCeTYTeyfe0-rN5r8zn9rw) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/streamlit@Z-inS5OKdXi6yveKgxMeC.md b/src/data/roadmaps/python-data-analysis/content/streamlit@Z-inS5OKdXi6yveKgxMeC.md index baa00ee47..46925b0b8 100644 --- a/src/data/roadmaps/python-data-analysis/content/streamlit@Z-inS5OKdXi6yveKgxMeC.md +++ b/src/data/roadmaps/python-data-analysis/content/streamlit@Z-inS5OKdXi6yveKgxMeC.md @@ -1,3 +1,9 @@ # Streamlit - -Streamlit is an open-source Python framework for building interactive data applications with minimal code. A Streamlit app is a Python script where each widget (slider, dropdown, text input) triggers a rerun of the script with the new value. It is popular for rapidly prototyping and sharing data analysis tools and ML demos. \ No newline at end of file + +Streamlit is an open-source Python framework for building interactive data applications with minimal code. A Streamlit app is a Python script where each widget (slider, dropdown, text input) triggers a rerun of the script with the new value. It is popular for rapidly prototyping and sharing data analysis tools and ML demos. + +Visit the following resources to learn more: + +- [@official@Tutorials](https://docs.streamlit.io/develop/tutorials) +- [@article@Getting Started with Streamlit](https://www.pythonguis.com/tutorials/getting-started-with-streamlit/) +- [@video@Streamlit: The Fastest Way To Build Python Apps?](https://www.youtube.com/watch?v=D0D4Pa22iG0) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/strings@F-0Rxwu_RcNF_NOoFHoo7.md b/src/data/roadmaps/python-data-analysis/content/strings@F-0Rxwu_RcNF_NOoFHoo7.md index c64ba8c85..22af7b19f 100644 --- a/src/data/roadmaps/python-data-analysis/content/strings@F-0Rxwu_RcNF_NOoFHoo7.md +++ b/src/data/roadmaps/python-data-analysis/content/strings@F-0Rxwu_RcNF_NOoFHoo7.md @@ -1,3 +1,8 @@ # Strings - -Strings (`str`) represent text data. Column names, categorical values, labels, and file paths are all strings. Python provides extensive string methods for cleaning, parsing, and transforming text data, and the `re` module adds regex-based pattern matching. \ No newline at end of file + +Strings (`str`) represent text data. Column names, categorical values, labels, and file paths are all strings. Python provides extensive string methods for cleaning, parsing, and transforming text data, and the `re` module adds regex-based pattern matching. + +Visit the following resources to learn more: + +- [@official@Built-in Types](https://docs.python.org/3/library/stdtypes.html) +- [@article@str](https://realpython.com/ref/builtin-types/str/) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/strip-replace-split@S1EzGF-CLu7tWDb1V6RQc.md b/src/data/roadmaps/python-data-analysis/content/strip-replace-split@S1EzGF-CLu7tWDb1V6RQc.md index ef5b591a9..743c49115 100644 --- a/src/data/roadmaps/python-data-analysis/content/strip-replace-split@S1EzGF-CLu7tWDb1V6RQc.md +++ b/src/data/roadmaps/python-data-analysis/content/strip-replace-split@S1EzGF-CLu7tWDb1V6RQc.md @@ -1,3 +1,8 @@ # strip, replace, split - -`strip()` removes leading and trailing whitespace from a string. `replace()` substitutes one substring with another. `split()` divides a string into a list based on a delimiter. These three methods are among the most used for basic text cleaning, applied either to Python strings directly or through the Pandas `.str` accessor on a column. \ No newline at end of file + +`strip()` removes leading and trailing whitespace from a string. `replace()` substitutes one substring with another. `split()` divides a string into a list based on a delimiter. These three methods are among the most used for basic text cleaning, applied either to Python strings directly or through the Pandas `.str` accessor on a column. + +Visit the following resources to learn more: + +- [@article@Python String Methods: split, join, replace, strip & More](https://devnook.dev/cheatsheets/python-string-methods-cheatsheet/) +- [@article@Python String Methods](https://www.w3schools.com/PYTHON/python_ref_string.asp) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/subplots-and-figures@GRg7HyeVuSwkKD5bCOFiY.md b/src/data/roadmaps/python-data-analysis/content/subplots-and-figures@GRg7HyeVuSwkKD5bCOFiY.md index 398fbff94..4304c596a 100644 --- a/src/data/roadmaps/python-data-analysis/content/subplots-and-figures@GRg7HyeVuSwkKD5bCOFiY.md +++ b/src/data/roadmaps/python-data-analysis/content/subplots-and-figures@GRg7HyeVuSwkKD5bCOFiY.md @@ -1,3 +1,9 @@ # Subplots and figures - -Matplotlib's `Figure` is the top-level container, and `Axes` objects are the individual plots within it. `plt.subplots(rows, cols)` creates a grid of axes for displaying multiple plots side by side. Subplots are used to compare distributions across groups or to show multiple related variables in one figure. \ No newline at end of file + +Matplotlib's `Figure` is the top-level container, and `Axes` objects are the individual plots within it. `plt.subplots(rows, cols)` creates a grid of axes for displaying multiple plots side by side. Subplots are used to compare distributions across groups or to show multiple related variables in one figure. + +Visit the following resources to learn more: + +- [@official@Create multiple subplots using plt.subplots](https://matplotlib.org/stable/gallery/subplots_axes_and_figures/subplots_demo.html) +- [@article@Matplotlib: Part 4. Subplots, Layouts, and Advanced Customizations](https://medium.com/@ebimsv/mastering-matplotlib-part-4-subplots-layouts-and-advanced-customizations-2f07c4a99e80) +- [@article@Introduction to Multi-figure Layouts](https://codesignal.com/learn/courses/customizing-and-styling-plots/lessons/multi-figure-layouts-with-matplotlib) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/tuples@olniprgaL7l-9YhwbFfmN.md b/src/data/roadmaps/python-data-analysis/content/tuples@olniprgaL7l-9YhwbFfmN.md index db9fcaf61..ab13adc05 100644 --- a/src/data/roadmaps/python-data-analysis/content/tuples@olniprgaL7l-9YhwbFfmN.md +++ b/src/data/roadmaps/python-data-analysis/content/tuples@olniprgaL7l-9YhwbFfmN.md @@ -1,3 +1,8 @@ # Tuples - -Tuples are ordered, immutable sequences. Unlike lists, their contents cannot be changed after creation. They are used to represent fixed collections of values, such as coordinate pairs, function return values, and dictionary keys where immutability is required. \ No newline at end of file + +Tuples are ordered, immutable sequences. Unlike lists, their contents cannot be changed after creation. They are used to represent fixed collections of values, such as coordinate pairs, function return values, and dictionary keys where immutability is required. + +Visit the following resources to learn more: + +- [@official@Tuple](https://docs.python.org/3/tutorial/datastructures.html#tuples-and-sequences) +- [@video@Python Tuples](https://www.youtube.com/watch?v=w6hL_dszMxk) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/type-casting@NjCor7ePiZapd4f6bMZlV.md b/src/data/roadmaps/python-data-analysis/content/type-casting@NjCor7ePiZapd4f6bMZlV.md index 5a7443188..44dc60822 100644 --- a/src/data/roadmaps/python-data-analysis/content/type-casting@NjCor7ePiZapd4f6bMZlV.md +++ b/src/data/roadmaps/python-data-analysis/content/type-casting@NjCor7ePiZapd4f6bMZlV.md @@ -1,3 +1,8 @@ # Type Casting - -Type casting converts a value from one data type to another using built-in functions like `int()`, `float()`, `str()`, and `bool()`. It is frequently needed when data is loaded with incorrect types, such as numbers stored as strings or booleans stored as integers. \ No newline at end of file + +Type casting converts a value from one data type to another using built-in functions like `int()`, `float()`, `str()`, and `bool()`. It is frequently needed when data is loaded with incorrect types, such as numbers stored as strings or booleans stored as integers. + +Visit the following resources to learn more: + +- [@article@Python Type Conversion](https://www.programiz.com/python-programming/type-conversion-and-casting) +- [@video@Learn type casting in 7 minutes!](https://www.youtube.com/watch?v=Qtq83lAoogM) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/variance--std-deviation@nZl1ngETFxKStt0tFvOpR.md b/src/data/roadmaps/python-data-analysis/content/variance--std-deviation@nZl1ngETFxKStt0tFvOpR.md index c667ed730..6756cf135 100644 --- a/src/data/roadmaps/python-data-analysis/content/variance--std-deviation@nZl1ngETFxKStt0tFvOpR.md +++ b/src/data/roadmaps/python-data-analysis/content/variance--std-deviation@nZl1ngETFxKStt0tFvOpR.md @@ -1,3 +1,8 @@ # Variance & Std. Deviation - -Variance measures the average squared deviation from the mean, and standard deviation is its square root in the original units. Both quantify how spread out data values are. Pandas computes them with `var()` and `std()`. A high standard deviation relative to the mean signals high variability in the data. \ No newline at end of file + +Variance measures the average squared deviation from the mean, and standard deviation is its square root in the original units. Both quantify how spread out data values are. Pandas computes them with `var()` and `std()`. A high standard deviation relative to the mean signals high variability in the data. + +Visit the following resources to learn more: + +- [@article@A Guide to Metrics in Exploratory Data Analysis](https://towardsdatascience.com/a-guide-to-metrics-in-exploratory-data-analysis-250b33f72297/) +- [@video@How to Calculate Standard Deviation & Variance in Python](https://www.youtube.com/watch?v=p4H2b2x_nWc) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/virtualenv--venv@IL0fFEs4CgK_eBft-SYAl.md b/src/data/roadmaps/python-data-analysis/content/virtualenv--venv@IL0fFEs4CgK_eBft-SYAl.md index 26d8c69e7..75da535e3 100644 --- a/src/data/roadmaps/python-data-analysis/content/virtualenv--venv@IL0fFEs4CgK_eBft-SYAl.md +++ b/src/data/roadmaps/python-data-analysis/content/virtualenv--venv@IL0fFEs4CgK_eBft-SYAl.md @@ -1,3 +1,9 @@ # virtualenv / venv - -`venv` is Python's built-in tool for creating isolated virtual environments. Each environment has its own Python interpreter and installed packages, preventing conflicts between projects. `virtualenv` is a third-party alternative with additional features. Virtual environments are a best practice for keeping project dependencies separate. \ No newline at end of file + +`venv` is Python's built-in tool for creating isolated virtual environments. Each environment has its own Python interpreter and installed packages, preventing conflicts between projects. `virtualenv` is a third-party alternative with additional features. Virtual environments are a best practice for keeping project dependencies separate. + +Visit the following resources to learn more: + +- [@official@venv β€” Creation of virtual environments](https://docs.python.org/3/library/venv.html) +- [@article@Python Virtual Environment](https://www.w3schools.com/python/python_virtualenv.asp) +- [@video@Python Virtual Environments - Full Tutorial for Beginners](https://www.youtube.com/watch?v=Y21OR1OPC9A) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/visual-inspection@UIFNQ9g5fMK-KoZCrqv27.md b/src/data/roadmaps/python-data-analysis/content/visual-inspection@UIFNQ9g5fMK-KoZCrqv27.md index 6570239b0..f5a0eb303 100644 --- a/src/data/roadmaps/python-data-analysis/content/visual-inspection@UIFNQ9g5fMK-KoZCrqv27.md +++ b/src/data/roadmaps/python-data-analysis/content/visual-inspection@UIFNQ9g5fMK-KoZCrqv27.md @@ -1,3 +1,8 @@ # Visual Inspection - -Visual inspection uses charts to identify outliers and anomalies that statistical thresholds might miss. Box plots show the IQR and flag points beyond the whiskers. Scatter plots reveal isolated points far from the main cluster. Histograms expose unusual spikes or gaps. Visual inspection is always a valuable complement to numerical outlier detection. \ No newline at end of file + +Visual inspection uses charts to identify outliers and anomalies that statistical thresholds might miss. Box plots show the IQR and flag points beyond the whiskers. Scatter plots reveal isolated points far from the main cluster. Histograms expose unusual spikes or gaps. Visual inspection is always a valuable complement to numerical outlier detection. + +Visit the following resources to learn more: + +- [@article@Using Boxplot To Identify Outliers For Continuous Variables](https://yashvaantlakham73.medium.com/26-pandas-data-cleaning-using-boxplot-to-identify-outliers-for-continuous-variables-6ace0e6023dd) +- [@article@5 Powerful Techniques with UDFs for Data Cleaning and 5 Graphical Methods](https://medium.com/@satyamdatasci/outlier-detection-in-python-5-powerful-techniques-with-udfs-for-data-cleaning-and-5-graphical-db053e7cf9bf) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/vs-code@ObA_xZDY7PxU54NGBwyVI.md b/src/data/roadmaps/python-data-analysis/content/vs-code@ObA_xZDY7PxU54NGBwyVI.md index a5385fd94..9b143302b 100644 --- a/src/data/roadmaps/python-data-analysis/content/vs-code@ObA_xZDY7PxU54NGBwyVI.md +++ b/src/data/roadmaps/python-data-analysis/content/vs-code@ObA_xZDY7PxU54NGBwyVI.md @@ -1,3 +1,8 @@ # VS Code - -Visual Studio Code is a lightweight, extensible code editor that supports Python development through extensions. With the Python and Jupyter extensions installed, VS Code supports notebooks, debugging, linting, and IntelliSense. It is a popular choice for analysts who want more IDE features than a browser-based notebook provides. \ No newline at end of file + +Visual Studio Code is a lightweight, extensible code editor that supports Python development through extensions. With the Python and Jupyter extensions installed, VS Code supports notebooks, debugging, linting, and IntelliSense. It is a popular choice for analysts who want more IDE features than a browser-based notebook provides. + +Visit the following resources to learn more: + +- [@official@Jupyter Notebooks in VS Code](https://code.visualstudio.com/docs/datascience/jupyter-notebooks) +- [@video@Getting Started with Jupyter Notebooks in VS Code](https://www.youtube.com/watch?v=suAkMeWJ1yE) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/working-with-strings@Sg5w8zO2Ji-uDJKEoWey9.md b/src/data/roadmaps/python-data-analysis/content/working-with-strings@Sg5w8zO2Ji-uDJKEoWey9.md index d95c5ee64..40dd0f0c6 100644 --- a/src/data/roadmaps/python-data-analysis/content/working-with-strings@Sg5w8zO2Ji-uDJKEoWey9.md +++ b/src/data/roadmaps/python-data-analysis/content/working-with-strings@Sg5w8zO2Ji-uDJKEoWey9.md @@ -1,3 +1,9 @@ # Working with Strings - -Python provides a rich set of string methods for manipulating text: `strip()`, `split()`, `replace()`, `upper()`, `lower()`, `startswith()`, `endswith()`, and many more. These methods are applied directly to strings or through Pandas' `.str` accessor for vectorized text cleaning on entire columns. \ No newline at end of file + +Python provides a rich set of string methods for manipulating text: `strip()`, `split()`, `replace()`, `upper()`, `lower()`, `startswith()`, `endswith()`, and many more. These methods are applied directly to strings or through Pandas' `.str` accessor for vectorized text cleaning on entire columns. + +Visit the following resources to learn more: + +- [@official@string β€” Common string operations](https://docs.python.org/3/library/string.html) +- [@video@String methods in Python are easy! 〰️](https://www.youtube.com/watch?v=tb6EYiHtcXU) +- [@video@Python Tutorial for Beginners 2: Strings](https://www.youtube.com/watch?v=k9TUPpGqYTo) \ No newline at end of file diff --git a/src/data/roadmaps/python-data-analysis/content/z-score@2HIeG6ywOA9BkE9W3Gu9v.md b/src/data/roadmaps/python-data-analysis/content/z-score@2HIeG6ywOA9BkE9W3Gu9v.md index d2c7cdcd6..4ce2571e5 100644 --- a/src/data/roadmaps/python-data-analysis/content/z-score@2HIeG6ywOA9BkE9W3Gu9v.md +++ b/src/data/roadmaps/python-data-analysis/content/z-score@2HIeG6ywOA9BkE9W3Gu9v.md @@ -1,3 +1,8 @@ # Z-score - -A Z-score measures how many standard deviations a value is from the mean. Values with a Z-score above 3 or below βˆ’3 are commonly flagged as outliers. Z-scores are calculated using `scipy.stats.zscore()` or manually and work best when the data is approximately normally distributed. \ No newline at end of file + +A Z-score measures how many standard deviations a value is from the mean. Values with a Z-score above 3 or below βˆ’3 are commonly flagged as outliers. Z-scores are calculated using `scipy.stats.zscore()` or manually and work best when the data is approximately normally distributed. + +Visit the following resources to learn more: + +- [@article@How to Find Outliers](https://codingnomads.com/how-to-find-outliers) +- [@video@Outlier detection and removal: z score, standard deviation](https://www.youtube.com/watch?v=KFuEAGR3HS4) \ No newline at end of file