diff --git a/README.md b/README.md index f9b1302..bdcaaf7 100644 --- a/README.md +++ b/README.md @@ -125,6 +125,7 @@ - [x] [DeepSeek-R1-Distill-Qwen-7B Langchain 接入](./models/DeepSeek-R1-Distill-Qwen/02-DeepSeek-R1-Distill-Qwen-7B%20Langchain%20接入.md) @骆秀韬 - [x] [DeepSeek-R1-Distill-Qwen-7B WebDemo 部署](./models/DeepSeek-R1-Distill-Qwen/03-DeepSeek-R1-Distill-Qwen-7B%20WebDemo%20部署.md) @骆秀韬 - [x] [DeepSeek-R1-Distill-Qwen-7B vLLM 部署调用](./models/DeepSeek-R1-Distill-Qwen/04-DeepSeek-R1-Distill-Qwen-7B%20vLLM%20部署调用.md) @骆秀韬 + - [x] [DeepSeek-R1-0528-Qwen3-8B-GRPO及swanlab可视化](./models/DeepSeek-R1-Distill-Qwen/05-DeepSeek-R1-0528-Qwen3-8B-GRPO及swanlab可视化.md) @郭宣伯 - [MiniCPM-o-2_6](https://github.com/OpenBMB/MiniCPM-o) - [x] [minicpm-o-2.6 FastApi 部署调用](./models/MiniCPM-o/01MiniCPM-o%202%206%20FastApi部署调用%20.md) @林恒宇 diff --git a/models/DeepSeek-R1-Distill-Qwen/05-DeepSeek-R1-0528-Qwen3-8B-GRPO及swanlab可视化.ipynb b/models/DeepSeek-R1-Distill-Qwen/05-DeepSeek-R1-0528-Qwen3-8B-GRPO及swanlab可视化.ipynb new file mode 100644 index 0000000..d162886 --- /dev/null +++ b/models/DeepSeek-R1-Distill-Qwen/05-DeepSeek-R1-0528-Qwen3-8B-GRPO及swanlab可视化.ipynb @@ -0,0 +1,29890 @@ +{ + "cells": [ + { + "cell_type": "code", + "execution_count": null, + "metadata": { + "id": "xsjuzuw-XphH" + }, + "outputs": [], + "source": [ + "# ==================================================\n", + "# 安装 unsloth 与指定版本 vllm 的依赖。\n", + "# ==================================================\n", + "\n", + "# pip install unsloth vllm==0.8.5.post1" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/" + }, + "id": "IuDt6gPSW_Ds", + "outputId": "a744116f-ffc9-4dc4-f47d-df35fa19217b" + }, + "outputs": [], + "source": [ + "# ==================================================\n", + "# 安装 langid 库,用于语言检测。\n", + "# ==================================================\n", + "\n", + "# pip install langid -qq" + ] + }, + { + "cell_type": "code", + "execution_count": 1, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/", + "height": 1000, + "referenced_widgets": [ + "ad395db5267344c4890ef02e5e4e1d4f", + "53cd752832ca45438a33b9e3fe8b0cd4", + "7c5656c2634543418d62009b469a7398", + "4452f2e00ebe44b687915136e85a7f89", + "5264d84060b541e3a3c6fe3ba3bfd8a8", + "5089a4f6e98748e1a7e7c3f1f3b43ae7", + "9c7c4cf7d8c6463c84f324220d089dff", + "2d54bb83ff5d4da7a4632a3d66a1373f", + "a4850b84abbf4cf1b56ba25a17f4dce2", + "1825f9b1db2046b1a6fd4123fd521317", + "ca439af43b9a4403aa72be5e13980450", + "0d54762ae2a24176b84f5d1b1d3afbbc", + "a2ba29de900642099294c6c56149e9b3", + "aaf72b0efc4941c9b48d0ac486cb1b37", + "c91a5c8f46a449e383398fbbb817c6e7", + "ec18ab0c53f24049aa4ff7fa216df498", + "c1b407fea8934069951ce8653435c427", + "7e28a3050b6a4b83aa2cdbc6d89128dc", + "be74b1ea62024e94b06a3c60cd83ef03", + "62cee809fc314419b6ee8a23b73ad9f4", + "d8ab91309cd14a8e8d5a9a98cd3bf6ac", + "a20385beff1b44678101760cbf2a9289", + "291ed9bc49bf423fbf542001368bf20d", + "eec61bdf29354de4a170a0b1cbb7f321", + "6a4af71abb6248909542c1baf00552ea", + "36fa4e90471d4fa7b4a2e320e64a7177", + "f7d94cb30e3040f592780b8b01fb03dc", + "d29cc3cae42c48bcba3b2f600f5723d6", + "57f3b440014947bfa60059ff565fe426", + "2d0472e60093417b8bf639c58fb63e5d", + "b41f9ce14e5d4bd686350d66434e9649", + "337f029ea24c4cce9bd4f373d1ff0065", + "7039cf21937d4fc0aad5e551f2de46b1", + "f6ac40a611b84dd9b0a96e8511ad43b8", + "1c8e497ba2ce4e01b50c55c56511afe3", + "5619359815e4404389994acd44172574", + "2f3db22fe43445e29512ee47ec31767f", + "ac1cd0df6fad41a587d54f5dff17ddba", + "df58eef8b3394b33be59bfff71cf39b2", + "babd6314d5a54d6e94e7ed0e562817f6", + "f4fbcebf2045446989fc2a4e16692023", + "3784857e231d4ddbba219f93401cfa8c", + "9e89da359da44b8b8ff7833cc1971fa6", + "fb9a8137dee24668a7bc014a483e2f90", + "2a784021410c42fc96ccd7ab2a930bfb", + "78a571587dc540bfb7ee64b3faae3dbc", + "40f970b0096d49a3ba83f2622ef993c4", + "00bcea27fb984aef9ca72390f2ccbd66", + "ef0796c6e52f4ef1bb7f3008e140e27d", + "5de1492b56874abcab6cacab85ffc3e2", + "eb832e8bd46744e380197d019c1c0a97", + "657ad4c6599046ecb768694b7fbac8b0", + "7f6e2e410b1d488cb57380312ad33062", + "635a7e8b526f4c9082f66c0a120f13b4", + "3ea6579b20a84744ad3aff349720329c", + "8651ad3e607542aabfe3036b6f617225", + "ce9e5bb7bacf4287898e533cb6b65671", + "8131234607aa41a3957dc3727dc058e5", + "315f991cb22e4e56b5c757d1b4305974", + "3d36dc33e0b546448482cbdea4bfd51d", + "d9ea4865bf7b435b8ec234efc884fda8", + "a66048eb5c3f41239070521a5fec2eba", + "32f51a8516144fa1aaed36a17c199931", + "deece2b54dc5434c8576b327e850618e", + "5fab61f28d484012bf8a554aa98077fc", + "ae37b07560c64ca8b2b58336f55d0b38", + "46362d5227cc4a4b833e8e0a3845f2e5", + "d375b1bd4ef24154b80a54a345410040", + "b16bcab6cf9f421f8309f8011ad30f12", + "06a12077a35f43d0a37c37cc4dfe48a8", + "6b489672d84e43df9defd5f587cc62c4", + "b7980c681c584b7781bf80dbb1c8f262", + "c8ddd5dfd05a4993bb7e4e5513070fc1", + "35a1f78ce0184218a539a4809f1eec97", + "652e6681592e4d72a6ef2a83bfffc8b6", + "dfb7ccc26bbc4841892458131b2cbc5d", + "f845df9e290040bcab51166539e2d0df", + "7bb2161a7ab44bc59e704d4fcbda5f83", + "8d7a46666ccc40c7817786c8f68f3613", + "e302cade968948dd89fea51e30d30e54", + "e783000239564e22a6a10b033a657408", + "88a34a4c3c444a81afb46f5887430532", + "d0a32668ce5b45fe9973f2fc3e9980d8", + "ac5a697f1ee345a3830e109b4de60508", + "1219e0209cbc497b9656988616799f46", + "bd222e300afb4ce592932bb69828bdb2", + "7c201ff993824bf6b0526bfaf344574c", + "78a1225925f042eabce9df1bcb352d5e", + "a849fe65867a4dd8a44a44158b36f156", + "73cf8bef5a944f11bcdb71e2180cd988", + "5ff7d924a91a45dfb73b8c5095cf06ac", + "d6df73ea477b4285967cdf71bd87e86b", + "3666b9de655b4c6b99b349ad20133098", + "f78a473dbaf44eb3ac496b001ae329d2", + "4d84285525c04defbb7eb9be22924537", + "deadaeca7c3c4e3ebdbdfb8059fd3f20", + "9ca7a471f185413aba80cdf826567ac2", + "7dce192f46b5482b8229ab998cb8553d", + "d23d3bed1041494b8807c6240002faff", + "0d16ec356b0d4a9fa9783935c1bf7e96", + "8e68dc40147c452783d5cf83c1408116", + "f63c7cdd63e5416db0f0922db3fd7908", + "d74a1e6c8355475ca070a7d578d89406", + "f208dd8658844152b7128e37a766da0e", + "99998ced16f048de877a86c3a150e43d", + "f20a935c9d9846fda2da9ffbbd302a6a", + "1516fa90d103489685e38eb97a3d6c44", + "aecccb9d09714ae2a878207f2ee45049", + "8f38456f9549478fba7474cb641d3985", + "91218d9042634f04ab54a1c3a0961bc6", + "84bb4652a47e4a61a60f0997e9dd3cc6", + "67342dc60a5d4ad8bdf5158a75e79e2e", + "c906225802b84aa48552280e73e741fc", + "666d00afc5c94f4c9500d81bff1e1457", + "a317980587ef4feb97d3eb2e0fd896f0", + "342dfbc3768049edadec012d61d097c1", + "e0ff47045a6c4ea78d8725b5e0972b61", + "6126329f463947b5a52d44f5a64d11a7", + "adf8d5bd602f4f53a608611b6fbec736", + "33e72fbc4c7e473ead8cab7511346a31", + "f0a98d994a2543e38ad0d927b44b82dd", + "ff264341a00d4a3eb6077413f82c164e", + "891ed1c3179b4bd9a4ed216ae6b2e51f", + "9abc3fa35cca4b889f0d325eb384c865", + "682b83bc3766438b84399efcdc866df9", + "15aa26a0829e4738b14a28f723dc192b", + "4a7f29671ac64829b8132b16503af35f", + "e6dd68e9ce7840c3818333f73b2bc674", + "5348e164820944c4bedd824b56bc6f81", + "bfb9ee43676c4162bf03d5b7a4aff097", + "f1f88e24679d4e1a90c86256f1c6dd2d", + "868129d4d4e04cddb654ada36d037e62", + "d01ca7d6c1ca4343b327055d5eeba673", + "d88fda96801d456bb8b6adb75cee10b2", + "a15356f8c0514e41a6e976896c2ae2d7", + "bf8052a4a83c4afc91e359fa06b9bf4f", + "78a3ba4b4cd0414a8a7723295e052023", + "6ab4305498cc4785ac27d4013bc419a2", + "e8cca5439ef24aabbc17cb04883ce84b", + "fa58952bf2a54a3fa1ee27088255b60a", + "cacd9c7a3ee748f496d1d28ac4069ead", + "96a5bb46c47b431ea479d3d04e3f9c8d", + "ffc7702cbb9b4c3d907b8c051c1322c3", + "16eaff0c30f5448cb89a63e04f86a3f8", + "8dc958d62a124f6a9f443cafee3aab20", + "ab163dc587cb479aa5c0c1e7acd48deb", + "b5b60cfde7d54b63a81c7cbba19dc9ed", + "39ba1e8385244ca4b86f148dbdf1a3c5", + "c4544a2669794cee8b6ea0f862dee5bd", + "3f605c4bd9b24aa7902e9cd1b20734d0", + "d8fa448f44ed43e98ab90cc686005723", + "098ddd74c04f4b478f4a5918a2a1c6f5", + "d4ec246b8ee1444d85e5e45417a074cc", + "176bed311f074b1db87b90f64f2e1c62", + "7124c93d897f455ea832d83a47e22398", + "b66b8205a1ac46968aabba1224b1b5d6", + "7f0923482366431bb5e742b367b85717", + "4f52adf53fd44910b7be69356998804b", + "dc73460dd339407a8ef70d70b0ff2908", + "18e1cf02ac8e496d9ec6c7043700bdee", + "caa9215388f844a0906556a35871aa92", + "39162ce93b574ee48f8d510748b023ff", + "9846b47582494a01ab0413a339b69ac2", + "0e31a330d9b441d2b8031d705e4525c6", + "116438a432534157854e55edd17290b0", + "68d5074e7b57464eade24db952ffa466", + "1b54692855d14197a412bd73a5d589d9", + "b002b5685ff84e7ea96836e778a27f31", + "3639b294cc494a2a84b997949fa4aa1c", + "bf202d6af8484ec88a247ca5d4d05a98", + "6799224cb66b493ba04ea79a8d11810c", + "4836517a4f53482d9d863e919f229939", + "19896a921c694bd49e6cd8e58243c9d5", + "46a39e03b09e4d4ba611be35a4a38a5f", + "397d2b8d165f4139911acd86d399da61", + "cd7cc23e780a472fabbd04579bd3fc79", + "6d6bb6b5551b4c4ea16546cdc9d55491", + "6d090dfa0a304b858c9fbd75875566a3", + "86ab8244e6d748acad911e7203bc5a2c", + "4cb5793862184ef38e2ad7637c750043", + "71c4b092ce184f2098c9b2d0b23e8923", + "acae05fde8de4c6daaede3de7b0b3a13", + "878fdcdafe4d4ec097c4b2313040fb72", + "d20d86afc1a94664b60b478fbce12f24", + "a4ec0efb0aac494f832c5b2669d2e1fc", + "7c045fd658ca416ea071ddb00f848bca", + "e7bb7f6b7e594b46ab95ea445df94196" + ] + }, + "id": "DkIvEkIIkEyB", + "outputId": "6a976e17-8336-4d93-840f-d9c0a588ca25" + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "🦥 Unsloth: Will patch your computer to enable 2x faster free finetuning.\n", + "🦥 Unsloth Zoo will now patch everything to make training faster!\n", + "INFO 07-03 15:42:45 [importing.py:53] Triton module has been replaced with a placeholder.\n", + "INFO 07-03 15:42:45 [__init__.py:239] Automatically detected platform cuda.\n" + ] + }, + { + "name": "stderr", + "output_type": "stream", + "text": [ + "Unrecognized keys in `rope_scaling` for 'rope_type'='yarn': {'attn_factor'}\n" + ] + }, + { + "name": "stdout", + "output_type": "stream", + "text": [ + "==((====))== Unsloth 2025.6.12: Fast Qwen3 patching. Transformers: 4.52.4. vLLM: 0.8.5.post1.\n", + " \\\\ /| NVIDIA A100-SXM4-80GB. Num GPUs = 1. Max memory: 79.151 GB. Platform: Linux.\n", + "O^O/ \\_/ \\ Torch: 2.6.0+cu124. CUDA: 8.0. CUDA Toolkit: 12.4. Triton: 3.2.0\n", + "\\ / Bfloat16 = TRUE. FA [Xformers = 0.0.29.post2. FA2 = True]\n", + " \"-____-\" Free license: http://github.com/unslothai/unsloth\n", + "Unsloth: Fast downloading is enabled - ignore downloading bars which are red colored!\n" + ] + }, + { + "name": "stderr", + "output_type": "stream", + "text": [ + "Unrecognized keys in `rope_scaling` for 'rope_type'='yarn': {'attn_factor'}\n", + "Unrecognized keys in `rope_scaling` for 'rope_type'='yarn': {'attn_factor'}\n" + ] + }, + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Unsloth: vLLM loading /opt/tiger/test0/DeepSeek-R1-0528-Qwen3-8B with actual GPU utilization = 69.6%\n", + "Unsloth: Your GPU has CUDA compute capability 8.0 with VRAM = 79.15 GB.\n", + "Unsloth: Using conservativeness = 1.0. Chunked prefill tokens = 1024. Num Sequences = 320.\n", + "Unsloth: vLLM's KV Cache can use up to 39.67 GB. Also swap space = 6 GB.\n", + "INFO 07-03 15:44:57 [config.py:717] This model supports multiple tasks: {'reward', 'score', 'generate', 'embed', 'classify'}. Defaulting to 'generate'.\n", + "INFO 07-03 15:44:57 [config.py:2003] Chunked prefill is enabled with max_num_batched_tokens=1024.\n", + "Unsloth: vLLM Bitsandbytes config using kwargs = {'load_in_8bit': False, 'load_in_4bit': True, 'bnb_4bit_compute_dtype': 'bfloat16', 'bnb_4bit_quant_storage': 'uint8', 'bnb_4bit_quant_type': 'fp4', 'bnb_4bit_use_double_quant': False, 'llm_int8_enable_fp32_cpu_offload': False, 'llm_int8_has_fp16_weight': False, 'llm_int8_skip_modules': [], 'llm_int8_threshold': 6.0}\n", + "INFO 07-03 15:44:58 [core.py:58] Initializing a V1 LLM engine (v0.8.5.post1) with config: model='/opt/tiger/test0/DeepSeek-R1-0528-Qwen3-8B', speculative_config=None, tokenizer='/opt/tiger/test0/DeepSeek-R1-0528-Qwen3-8B', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, override_neuron_config=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=1024, download_dir=None, load_format=LoadFormat.BITSANDBYTES, tensor_parallel_size=1, pipeline_parallel_size=1, disable_custom_all_reduce=False, quantization=bitsandbytes, enforce_eager=False, kv_cache_dtype=auto, device_config=cuda:0, decoding_config=DecodingConfig(guided_decoding_backend='auto', reasoning_backend=None), observability_config=ObservabilityConfig(show_hidden_metrics=False, otlp_traces_endpoint=None, collect_model_forward_time=False, collect_model_execute_time=False), seed=0, served_model_name=/opt/tiger/test0/DeepSeek-R1-0528-Qwen3-8B, num_scheduler_steps=1, multi_step_stream_outputs=True, enable_prefix_caching=True, chunked_prefill_enabled=True, use_async_output_proc=True, disable_mm_preprocessor_cache=False, mm_processor_kwargs=None, pooler_config=None, compilation_config={\"level\":3,\"backend\":\"inductor\",\"custom_ops\":[\"none\"],\"splitting_ops\":[\"vllm.unified_attention\",\"vllm.unified_attention_with_output\"],\"use_inductor\":true,\"compile_sizes\":[],\"inductor_compile_config\":{\"debug\":false,\"dce\":true,\"coordinate_descent_tuning\":true,\"trace.enabled\":false,\"trace.graph_diagram\":false,\"triton.cudagraphs\":true,\"compile_threads\":48,\"max_autotune\":false,\"disable_progress\":false,\"verbose_progress\":true,\"enable_auto_functionalized_v2\":false},\"use_cudagraph\":true,\"cudagraph_num_of_warmups\":1,\"cudagraph_capture_sizes\":[512,504,496,488,480,472,464,456,448,440,432,424,416,408,400,392,384,376,368,360,352,344,336,328,320,312,304,296,288,280,272,264,256,248,240,232,224,216,208,200,192,184,176,168,160,152,144,136,128,120,112,104,96,88,80,72,64,56,48,40,32,24,16,8,4,2,1],\"max_capture_size\":512}\n", + "WARNING 07-03 15:44:58 [utils.py:2522] Methods determine_num_available_blocks,device_config,get_cache_block_size_bytes,initialize_cache not implemented in \n", + "INFO 07-03 15:44:58 [parallel_state.py:1004] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, TP rank 0\n", + "INFO 07-03 15:44:58 [cuda.py:221] Using Flash Attention backend on V1 engine.\n", + "WARNING 07-03 15:44:58 [topk_topp_sampler.py:69] FlashInfer is not available. Falling back to the PyTorch-native implementation of top-p & top-k sampling. For the best performance, please install FlashInfer.\n", + "INFO 07-03 15:44:58 [gpu_model_runner.py:1329] Starting to load model /opt/tiger/test0/DeepSeek-R1-0528-Qwen3-8B...\n", + "INFO 07-03 15:44:58 [loader.py:1187] Loading weights with BitsAndBytes quantization. May take a while ...\n" + ] + }, + { + "data": { + "application/vnd.jupyter.widget-view+json": { + "model_id": "4556ee1c3b07499ba594c9ee317d128a", + "version_major": 2, + "version_minor": 0 + }, + "text/plain": [ + "Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<|User|>What is 1+1?<|Assistant|>2<|end▁of▁sentence|><|User|>What is 1+1?<|Assistant|>2<|end▁of▁sentence|><|Assistant|>\n" + ] + } + ], + "source": [ + "# ==================================================\n", + "# 演示使用 tokenizer.apply_chat_template 将多轮消息转换为模型可用的 token。\n", + "# ==================================================\n", + "\n", + "print(tokenizer.apply_chat_template([\n", + " {\"role\" : \"user\", \"content\" : \"What is 1+1?\"},\n", + " {\"role\" : \"assistant\", \"content\" : f\"I think it's 2.22\"},\n", + " {\"role\" : \"user\", \"content\" : \"What is 1+1?\"},\n", + " {\"role\" : \"assistant\", \"content\" : f\"I think it's 2.22\"},\n", + "], tokenize = False, add_generation_prompt = True))" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "id": "7KGgPgk_5S8r" + }, + "source": [ + "### 数据准备\n", + "\n", + "\n", + "我们将使用 Hugging Face 的 [Open R1 Math 数据集](https://huggingface.co/datasets/openr1) 以及 [GSM8K 数据集](https://huggingface.co/datasets/openai/gsm8k)。" + ] + }, + { + "cell_type": "code", + "execution_count": 4, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/", + "height": 201, + "referenced_widgets": [ + "c5560a8612a34eaf96630cd90454cb2d", + "8ab97b85d9484a6cbf654cbc954f75ff", + "4b6aac4b3426447bbe45fdd14bcce426", + "58ef324089874a49a1fe7e5f94ff93f0", + "bbce5a3db2384bcc8ebe20d736aacd2d", + "ee569d9509e54678a321902fdf9a040f", + "fe1d739b0b5941bd9d23bf137b0e98ed", + "839d4de9dc6c403d8db455c37093b5b5", + "756b6715039c4c75b38f9f711eca51c8", + "73a15eb988404223b07a8e3233a1d549", + "db197b2558d946608af02e62751c1de9", + "dd08b63a04f6409987863cd204d64796", + "f49189f73d4e4031a086b27eb3e54e4f", + "0f4cb21a4bb84aafa59ef9baa8511714", + "e16262d41d5247b98e5c946f04e2b8b1", + "e1feda9d8d22452e9d253b6635af7caa", + "d479e0e9318547099bb549e32897481a", + "aa51b6a90f304328840503bc39898320", + "ee904e35b4654c6a8ca8a9bd47924a86", + "1b857d475d0d400d9ac2aaacc55cfd84", + "a6f28bc3c2a440efbae14dab1246de1c", + "325a955756a549ffb375150dfd6c058a", + "4db31459944a4e84b1335981f70067f5", + "e9a22bd6ca034cb89c2e218eb63ab520", + "67c8b33fbc9e482a875e02bb7750916e", + "82f6146a5ac847a58a4d9ea42756464f", + "57bda5a2a1fd4366b2cd454f055c9d5a", + "f48756b6ba4d4189a16d4092bf936677", + "fa1cf861a0ec4ba19e0dc72fb1a9edd4", + "4d97073b9a1e4071a13dd95c77efe4cc", + "eefea66fa97d45e9a007ae77f03dcb1d", + "b6fad7ea371b4ef798f150915623b72c", + "43f18a400411445abe16be9b174c5726" + ] + }, + "id": "o7-eUrQn-OzE", + "outputId": "a16658cc-4b2b-4c55-d649-c37cda2dada5" + }, + "outputs": [ + { + "data": { + "application/vnd.jupyter.widget-view+json": { + "model_id": "f8efb120aadf46a38107d3836858d9b8", + "version_major": 2, + "version_minor": 0 + }, + "text/plain": [ + "Generating train split: 0%| | 0/14116 [00:00(.*)', re.DOTALL|re.UNICODE)" + ] + }, + "execution_count": 9, + "metadata": {}, + "output_type": "execute_result" + } + ], + "source": [ + "# ==================================================\n", + "# 引入正则表达式,定义 solution_end_regex 以及辅助语言检测函数。\n", + "# ==================================================\n", + "\n", + "import re\n", + "\n", + "# Add optional EOS token matching\n", + "solution_end_regex = rf\"{reasoning_end}(.*)\"\n", + "\n", + "match_format = re.compile(solution_end_regex, re.DOTALL)\n", + "match_format" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "id": "OycMneOq-iNC" + }, + "source": [ + "验证正则表达式是否有效:" + ] + }, + { + "cell_type": "code", + "execution_count": 10, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/" + }, + "id": "ndzHnQ_6-jHt", + "outputId": "f81976c2-45fb-4fdb-c02e-86297592a43c" + }, + "outputs": [ + { + "data": { + "text/plain": [ + "['Hence, the solution is 2.']" + ] + }, + "execution_count": 10, + "metadata": {}, + "output_type": "execute_result" + } + ], + "source": [ + "# ==================================================\n", + "# 使用 match_format 正则,对示例字符串进行格式匹配测试(未加 标签)。\n", + "# ==================================================\n", + "\n", + "match_format.findall(\n", + " \"Let me think!\"\\\n", + " f\"Hence, the solution is 2.\",\n", + ")" + ] + }, + { + "cell_type": "code", + "execution_count": 11, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/" + }, + "id": "eRMDAzDk2x6t", + "outputId": "593d7596-8b7f-4769-eca9-bbb6b15c0515" + }, + "outputs": [ + { + "data": { + "text/plain": [ + "['\\n\\nHence, the solution is 2']" + ] + }, + "execution_count": 11, + "metadata": {}, + "output_type": "execute_result" + } + ], + "source": [ + "# ==================================================\n", + "# 使用 match_format 正则,对示例字符串进行格式匹配测试(含 标签)。\n", + "# ==================================================\n", + "\n", + "match_format.findall(\n", + " \"Let me think!\"\\\n", + " f\"\\n\\nHence, the solution is 2\",\n", + ")" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "id": "weOjmO5l-kl3" + }, + "source": [ + "接下来我们要创建一个奖励函数:若输出格式完全正确则给 3 分:" + ] + }, + { + "cell_type": "code", + "execution_count": 12, + "metadata": { + "id": "qgFNXORy-lpO" + }, + "outputs": [], + "source": [ + "# ==================================================\n", + "# 定义 match_format_exactly 函数:若返回内容严格符合模板格式,则给予高奖励。\n", + "# ==================================================\n", + "\n", + "def match_format_exactly(completions, **kwargs):\n", + " scores = []\n", + " for completion in completions:\n", + " score = 0\n", + " response = completion[0][\"content\"]\n", + " if match_format.search(response) is not None: score += 3.0\n", + " scores.append(score)\n", + " return scores" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "id": "Gf69i2WT-m4K" + }, + "source": [ + "若完全匹配失败,则按符号数量部分奖励模型:" + ] + }, + { + "cell_type": "code", + "execution_count": 13, + "metadata": { + "id": "cUfHzCVx-nGK" + }, + "outputs": [], + "source": [ + "# ==================================================\n", + "# 定义 match_format_approximately 函数:当模板部分符合时,以符号数量作为次级奖励。\n", + "# ==================================================\n", + "\n", + "def match_format_approximately(completions, **kwargs):\n", + " scores = []\n", + " for completion in completions:\n", + " score = 0\n", + " response = completion[0][\"content\"]\n", + " score += 0.5 if response.count(reasoning_start) == 1 else -1.0\n", + " score += 0.5 if response.count(reasoning_end) == 1 else -1.0\n", + " scores.append(score)\n", + " return scores" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "id": "9wAUWwtE-s6n" + }, + "source": [ + "提取模型生成的答案,根据其与真实答案的接近程度按比值奖励或惩罚:" + ] + }, + { + "cell_type": "code", + "execution_count": 14, + "metadata": { + "id": "hmtI_8gg-uIE" + }, + "outputs": [], + "source": [ + "# ==================================================\n", + "# 定义 check_answer 函数:从预测结果里抽取答案,与真值比较并计算分数。\n", + "# ==================================================\n", + "\n", + "def check_answer(prompts, completions, answer, **kwargs):\n", + " question = prompts[0][-1][\"content\"]\n", + " responses = [completion[0][\"content\"] for completion in completions]\n", + "\n", + " extracted_responses = [\n", + " guess.group(1)\n", + " if (guess := match_format.search(r)) is not None else None \\\n", + " for r in responses\n", + " ]\n", + "\n", + " scores = []\n", + " for guess, true_answer in zip(extracted_responses, answer):\n", + " score = 0\n", + " if guess is None:\n", + " scores.append(-2.0)\n", + " continue\n", + " # Correct answer gets 5 points!\n", + " if guess == true_answer:\n", + " score += 5.0\n", + " # Match if spaces are seen, but less reward\n", + " elif guess.strip() == true_answer.strip():\n", + " score += 3.5\n", + " else:\n", + " # We also reward it if the answer is close via ratios!\n", + " # Ie if the answer is within some range, reward it!\n", + " try:\n", + " ratio = float(guess) / float(true_answer)\n", + " if ratio >= 0.9 and ratio <= 1.1: score += 2.0\n", + " elif ratio >= 0.8 and ratio <= 1.2: score += 1.5\n", + " else: score -= 2.5 # Penalize wrong answers\n", + " except:\n", + " score -= 4.5 # Penalize\n", + " scores.append(score)\n", + " return scores" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "id": "atMyfhXh-v3R" + }, + "source": [ + "有时答案不止一个数字,例如 20。\n", + "\n", + "我们还会移除类似 123,456 中的逗号分隔符。" + ] + }, + { + "cell_type": "code", + "execution_count": 15, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/" + }, + "id": "AVW0kL8q-wL5", + "outputId": "0bb3b3ab-47ae-4d59-c1c1-2c20a3a9efff" + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "['0.34']\n", + "['123,456']\n", + "['-0.234']\n", + "['17']\n" + ] + } + ], + "source": [ + "# ==================================================\n", + "# 测试 match_numbers 正则表达式在不同数字字符串上的提取效果。\n", + "# ==================================================\n", + "\n", + "match_numbers = re.compile(\n", + " r\".*?[\\s]{0,}([-]?[\\d\\.\\,]{1,})\",\n", + " flags = re.MULTILINE | re.DOTALL\n", + ")\n", + "print(match_numbers.findall(\" 0.34 \"))\n", + "print(match_numbers.findall(\" 123,456 \"))\n", + "print(match_numbers.findall(\" -0.234 \"))\n", + "print(match_numbers.findall(\"17\"))" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "id": "19KD28CXW_EO" + }, + "source": [ + "最后,我们引入 DeepSeek R1 论文中的 `language consistency reward`,以确保思考过程保持一致的目标语言。" + ] + }, + { + "cell_type": "code", + "execution_count": 16, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/" + }, + "id": "1PU4QSGHW_EO", + "outputId": "be13cfd3-fbd1-43b6-8567-5ac06e020038" + }, + "outputs": [ + { + "name": "stderr", + "output_type": "stream", + "text": [ + "[2025-07-03 15:54:32] INFO langid.py:162: initializing identifier\n" + ] + }, + { + "name": "stdout", + "output_type": "stream", + "text": [ + "en\n", + "id\n", + "zh\n" + ] + } + ], + "source": [ + "# ==================================================\n", + "# 使用 langid 库定义 get_lang 函数,检测输入文本的语言类别。\n", + "# ==================================================\n", + "\n", + "import langid\n", + "\n", + "def get_lang(text: str) -> str:\n", + " if not text:\n", + " return \"und\"\n", + " lang, _ = langid.classify(text)\n", + " return lang\n", + "\n", + "\n", + "print(get_lang(\"Hello, How are you\")) # This should return en\n", + "print(get_lang(\"Aku berpikir kalau aku adalah kamu\")) # This should return id\n", + "print(get_lang(\"我在这里\")) # This should return zh" + ] + }, + { + "cell_type": "code", + "execution_count": 17, + "metadata": { + "id": "czn2loIDW_EQ" + }, + "outputs": [], + "source": [ + "# ==================================================\n", + "# 定义 format_and_language_reward_func:综合格式匹配与目标语言一致性进行奖励。\n", + "# ==================================================\n", + "\n", + "import re\n", + "\n", + "def format_and_language_reward_func(completions, **kwargs):\n", + " scores = []\n", + "\n", + " for completion_item in completions:\n", + " if not completion_item or not isinstance(completion_item[0], dict) or \"content\" not in completion_item[0]:\n", + " scores.append(-5.0)\n", + " print(f\"Warning: Malformed completion item, assigning default low score: {completion_item}\")\n", + " continue\n", + "\n", + " content = completion_item[0][\"content\"]\n", + "\n", + " lang = get_lang(content)\n", + "\n", + " if lang == 'id':\n", + " score = 5.0\n", + " elif lang == 'en':\n", + " score = -3.0\n", + " elif lang == 'zh':\n", + " score = -3.0\n", + " else:\n", + " score = -5.0\n", + "\n", + " scores.append(score)\n", + "\n", + " return scores" + ] + }, + { + "cell_type": "code", + "execution_count": 18, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/" + }, + "id": "XjTfmkTAW_ER", + "outputId": "4e70daa7-3ac0-4e66-ead5-8541b7905185" + }, + "outputs": [ + { + "data": { + "text/plain": [ + "[-3.0, -3.0]" + ] + }, + "execution_count": 18, + "metadata": {}, + "output_type": "execute_result" + } + ], + "source": [ + "# ==================================================\n", + "# 示例:调用 format_and_language_reward_func 并打印得分结果。\n", + "# ==================================================\n", + "\n", + "prompts = [\n", + " [{\"role\": \"assistant\", \"content\": \"What is the result of (1 + 2) * 4?\"}],\n", + " [{\"role\": \"assistant\", \"content\": \"What is the result of (3 + 1) * 2?\"}],\n", + "]\n", + "completions = [\n", + " [{\"role\": \"assistant\", \"content\": \"The sum of 1 and 2 is 3, which we multiply by 4 to get 12.(1 + 2) * 4 = 12\"}],\n", + " [{\"role\": \"assistant\", \"content\": \"The sum of 3 and 1 is 4, which we multiply by 2 to get 8. So (3 + 1) * 2 = 8.\"}],\n", + "]\n", + "format_and_language_reward_func(prompts=prompts, completions=completions)" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "id": "RbfaaAywNHHh" + }, + "source": [ + "准备主函数:打印生成答案,并将文本转换为浮点数以比较是否正确。" + ] + }, + { + "cell_type": "code", + "execution_count": 19, + "metadata": { + "id": "GjBFrttr-y1_" + }, + "outputs": [], + "source": [ + "# ==================================================\n", + "# 定义 check_numbers 函数:比较预测答案与真实数字之间的误差,并打印日志。\n", + "# ==================================================\n", + "\n", + "global PRINTED_TIMES\n", + "PRINTED_TIMES = 0\n", + "global PRINT_EVERY_STEPS\n", + "PRINT_EVERY_STEPS = 5\n", + "\n", + "def check_numbers(prompts, completions, answer, **kwargs):\n", + " question = prompts[0][-1][\"content\"]\n", + " responses = [completion[0][\"content\"] for completion in completions]\n", + "\n", + " extracted_responses = [\n", + " guess.group(1)\n", + " if (guess := match_numbers.search(r)) is not None else None \\\n", + " for r in responses\n", + " ]\n", + "\n", + " scores = []\n", + " # Print only every few steps\n", + " global PRINTED_TIMES\n", + " global PRINT_EVERY_STEPS\n", + " if PRINTED_TIMES % PRINT_EVERY_STEPS == 0:\n", + " print(\n", + " '*'*20 + f\"Question:\\n{question}\", f\"\\nAnswer:\\n{answer[0]}\", f\"\\nResponse:\\n{responses[0]}\", f\"\\nExtracted:\\n{extracted_responses[0]}\"\n", + " )\n", + " PRINTED_TIMES += 1\n", + "\n", + " for guess, true_answer in zip(extracted_responses, answer):\n", + " if guess is None:\n", + " scores.append(-2.5)\n", + " continue\n", + " # Convert to numbers\n", + " try:\n", + " true_answer = float(true_answer.strip())\n", + " # Remove commas like in 123,456\n", + " guess = float(guess.strip().replace(\",\", \"\"))\n", + " scores.append(3.5 if guess == true_answer else -1.5)\n", + " except:\n", + " scores.append(0)\n", + " continue\n", + " return scores" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "id": "fgOR3wJ_AyLr" + }, + "source": [ + "获取提示长度的 90% 分位数,避免过长的提示被截断(删除最前 10% 的超长提示)。" + ] + }, + { + "cell_type": "code", + "execution_count": 20, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/", + "height": 171, + "referenced_widgets": [ + "efe86cb2e7174a149f1c01544b1f9d4f", + "73b5b814482f4d1faa3cbd8d59520b19", + "1e53f50fc9de48f48919b3eef416d825", + "c0b6b9c6c70946f292cbd3ef194647f3", + "2c4bbe0814524b8dae16f0ea435f388d", + "4d9c429efca2406ca4dbb606fd888f27", + "2310e1c0d376470990d063de29eabb68", + "a58bf1ccd1614069b87c278b14620636", + "27c53c1f5c934e1eb381ed302f186b85", + "b8d562a3b0fb40e0a15d87830f90aa4c", + "b2520312d2a345b3913ef02bca7d060f", + "4bdd6660598447aab9bd8249d98ec407", + "6bdca654291e410da6f0859be242a3cd", + "38b11c99e03c454499f6013ed0326e4c", + "74c2df2966ab4e30b27a326a6f39993b", + "92066c72faba4956a29bde4deaa0fed3", + "a670d8e6c3414a35a369476b77a178ac", + "976c8cfbb4cd444a83f28a65c75e057d", + "efff7b47e0c543c1b579ca9f855f4954", + "1eff3897c2af43fd8be2d2cd0bc7b512", + "57d307bad8e7431eb9d9562ffafd7d17", + "f77640ae66ae410d9a1858cf7880a192" + ] + }, + "id": "6EgAi4Q5fGE-", + "outputId": "e3719f82-3c6a-4cb3-9bc1-995cd67689a9" + }, + "outputs": [ + { + "data": { + "application/vnd.jupyter.widget-view+json": { + "model_id": "8ecf0759a0eb4c26a81c4d1212c9f78f", + "version_major": 2, + "version_minor": 0 + }, + "text/plain": [ + "Map: 0%| | 0/14116 [00:00You are given a problem.\n", + "Think about the problem and provide your working out.\n", + "You must think in Bahasa Indonesia.<|User|>In triangle $ABC$, $\\sin \\angle A = \\frac{4}{5}$ and $\\angle A < 90^\\circ$. Let $D$ be a point outside triangle $ABC$ such that $\\angle BAD = \\angle DAC$ and $\\angle BDC = 90^\\circ$. Suppose that $AD = 1$ and that $\\frac{BD}{CD} = \\frac{3}{2}$. If $AB + AC$ can be expressed in the form $\\frac{a\\sqrt{b}}{c}$ where $a, b, c$ are pairwise relatively prime integers, find $a + b + c$.<|Assistant|>\n" + ] + }, + { + "data": { + "application/vnd.jupyter.widget-view+json": { + "model_id": "da855afe07b14c29ac5b8622baceb149", + "version_major": 2, + "version_minor": 0 + }, + "text/plain": [ + "Map: 0%| | 0/14116 [00:00\n", + "### 训练模型\n", + "\n", + "现在配置 GRPO Trainer 及所有参数!" + ] + }, + { + "cell_type": "code", + "execution_count": 21, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/" + }, + "id": "ptqkXK2D4d6p", + "outputId": "bbabb969-329e-48c3-d92e-c27f35fc7766" + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Unsloth: We now expect `per_device_train_batch_size` to be a multiple of `num_generations`.\n", + "We will change the batch size of 1 to the `num_generations` of 4\n" + ] + } + ], + "source": [ + "# ==================================================\n", + "# 设置 vLLM 的采样参数,例如 temperature、top_k、最大生成长度等。\n", + "# ==================================================\n", + "\n", + "max_prompt_length = maximum_length + 1 # + 1 just in case!\n", + "max_completion_length = max_seq_length - max_prompt_length\n", + "\n", + "from vllm import SamplingParams\n", + "vllm_sampling_params = SamplingParams(\n", + " min_p = 0.1,\n", + " top_p = 1.0,\n", + " top_k = -1,\n", + " seed = 3407,\n", + " stop = [tokenizer.eos_token],\n", + " include_stop_str_in_output = True,\n", + ")\n", + "\n", + "from trl import GRPOConfig, GRPOTrainer\n", + "training_args = GRPOConfig(\n", + " vllm_sampling_params = vllm_sampling_params,\n", + " temperature = 1.0,\n", + " learning_rate = 5e-6,\n", + " weight_decay = 0.01,\n", + " warmup_ratio = 0.1,\n", + " lr_scheduler_type = \"linear\",\n", + " optim = \"adamw_8bit\",\n", + " logging_steps = 1,\n", + " per_device_train_batch_size = 1,\n", + " gradient_accumulation_steps = 1, # Increase to 4 for smoother training\n", + " num_generations = 4, # Decrease if out of memory\n", + " max_prompt_length = max_prompt_length,\n", + " max_completion_length = max_completion_length,\n", + " # num_train_epochs = 1, # Set to 1 for a full training run\n", + " max_steps = 100,\n", + " save_steps = 100,\n", + " report_to = \"swanlab\", # Can use Weights & Biases\n", + " output_dir = \"outputs\",\n", + "\n", + " # For optional training + evaluation\n", + " # fp16_full_eval = True,\n", + " # per_device_eval_batch_size = 4,\n", + " # eval_accumulation_steps = 1,\n", + " # eval_strategy = \"steps\",\n", + " # eval_steps = 1,\n", + ")" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "id": "r9Mv8UZO5hz-" + }, + "source": [ + "开始运行 Trainer!向上滚动可查看训练过程中生成的表格日志。" + ] + }, + { + "cell_type": "code", + "execution_count": 22, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/", + "height": 1000 + }, + "id": "vzOuSVCL_GA9", + "outputId": "58162095-173b-465e-af98-3c717e8d8424" + }, + "outputs": [ + { + "name": "stderr", + "output_type": "stream", + "text": [ + "[2025-07-03 15:56:05] WARNING other.py:492: Detected kernel version 5.4.143, which is below the recommended minimum of 5.5.0; this can cause the process to hang. It is recommended to upgrade the kernel to the minimum version or higher.\n", + "==((====))== Unsloth - 2x faster free finetuning | Num GPUs used = 1\n", + " \\\\ /| Num examples = 12,728 | Num Epochs = 1 | Total steps = 100\n", + "O^O/ \\_/ \\ Batch size per device = 4 | Gradient accumulation steps = 1\n", + "\\ / Data Parallel GPUs = 1 | Total batch size (4 x 1 x 1) = 4\n", + " \"-____-\" Trainable parameters = 87,293,952 of 8,000,000,000 (1.09% trained)\n" + ] + }, + { + "name": "stdout", + "output_type": "stream", + "text": [ + "\u001b[1m\u001b[34mswanlab\u001b[0m\u001b[0m: Tracking run with swanlab version 0.6.4 \n", + "\u001b[1m\u001b[34mswanlab\u001b[0m\u001b[0m: Run data will be saved locally in \u001b[35m\u001b[1m/opt/tiger/test0/swanlog/run-20250703_155610-0e8cd89d\u001b[0m\u001b[0m\n", + "\u001b[1m\u001b[34mswanlab\u001b[0m\u001b[0m: 👋 Hi \u001b[1m\u001b[39mtwosugar\u001b[0m\u001b[0m, welcome to swanlab!\n", + "\u001b[1m\u001b[34mswanlab\u001b[0m\u001b[0m: Syncing run \u001b[33moutputs\u001b[0m to the cloud\n", + "\u001b[1m\u001b[34mswanlab\u001b[0m\u001b[0m: 🏠 View project at \u001b[34m\u001b[4mhttps://swanlab.cn/@twosugar/test0\u001b[0m\u001b[0m\n", + "\u001b[1m\u001b[34mswanlab\u001b[0m\u001b[0m: 🚀 View run at \u001b[34m\u001b[4mhttps://swanlab.cn/@twosugar/test0/runs/veyt1tdlqckdz81hdyqmc\u001b[0m\u001b[0m\n" + ] + }, + { + "data": { + "text/html": [ + "\n", + "\n", + "\n", + "\n", + " \n", + " \n", + " Show Iframe\n", + " \n", + " \n", + " \n", + "\n", + "\n", + "

" + ], + "text/plain": [ + "" + ] + }, + "metadata": {}, + "output_type": "display_data" + }, + { + "name": "stdout", + "output_type": "stream", + "text": [ + "********************Question:\n", + "In the diagram, each of the three identical circles touch the other two. The circumference of each circle is 36. What is the perimeter of the shaded region? [asy]\n", + "\n", + "defaultpen(1);\n", + "\n", + "path p = (1, 0){down}..{-dir(30)}dir(-60){dir(30)}..{dir(-30)}((2, 0) + dir(-120)){-dir(-30)}..{up}(1, 0)--cycle;\n", + "fill(p, gray(0.75));\n", + "\n", + "draw(unitcircle);\n", + "draw(shift(2 * dir(-60)) * unitcircle);\n", + "draw(shift(2) * unitcircle);\n", + "[/asy] \n", + "Answer:\n", + "18 \n", + "Response:\n", + "\n", + "Pertama, soalnya: Ada tiga lingkaran identik yang saling menyentuh. Keliling setiap lingkaran adalah 36. We're asked for the perimeter of the shaded region, dan dari Asymptote diagram, tampaknya shaded region adalah daerah yang diisi oleh kurva ketiga lingkaran.\n", + "\n", + "Jadi, ada tiga lingkaran yang sama besar, dan masing-masing menyentuh lingkaran lainnya. Ini menimbulkan pertanyaan: di mana persentuhan ini terjadi? Dalam diagram Asymptote, tampaknya lingkaran berada di posisi yang membentuk suatu pola segitiga. Langsung saja, karena keliling diketahui, kita bisa mencari jari-jari.\n", + "\n", + "Keliling lingkaran C = 2 * π * r = 36. Jadi, untuk satu lingkaran, r = 36 / (2 * π) = 18 / π. Tapi kita mungkin perlu menggunakan π dalam perhitungan, atau mungkin semuanya akan sederhana nanti.\n", + "\n", + "Sekarang, untuk mencari perimeter of shaded region. Dari diagram, shaded region sepertinya adalah daerah yang \"bervisdom\" atau daerah yang dibatasi oleh tiga busur lingkaran.\n", + "\n", + "Dalam Asymptote code, tertulis:\n", + "\n", + "path p = (1, 0){down}..{-dir(30)}dir(-60){dir(30)}..{dir(-30)}((2, 0) + dir(-120)){-dir(-30)}..{up}(1, 0)--cycle;\n", + "\n", + " It looks messy, but I think it's defining a path with Bezier curves. From the draw commands:\n", + "\n", + "draw(unitcircle);\n", + "draw(shift(2 * dir(-60)) * unitcircle);\n", + "draw(shift(2) * unitcircle);\n", + "\n", + "unitcircle is probably not the same as the circles with circumference 36, but in the context, likely it's meant to be unit circle, but unit circle is radius 1? Probably not, because the code should match the actual circles.\n", + "\n", + "The \"unitcircle\" might be a placeholder, but in the path p, it's using dir(30) etc, which are angles.\n", + "\n", + "Since the circles are identical and each other touching, the shaded region is likely the region enclosed by the three circles, which is a Reuleaux triangle or something, but Reuleaux is made from three circles.\n", + "\n", + "A Reuleaux triangle is formed by taking three circles, each centered at the other two vertices.\n", + "\n", + "In this case, with three circles each touching the other two.\n", + "\n", + "I need to find the perimeter of the shaded region. From the Asympt, it seems that the shaded region is traced with a path that goes through three points and is defined with Bezier curves, but likely it means the boundary of the three circles that are not overlapping or something.\n", + "\n", + "Since all circles have circumference 36, radius r = 36 / (2π) = 18/π.\n", + "\n", + "Now, the key point is how these circles are arranged. If each circle touches the other two, then the distance between centers should be 2r.\n", + "\n", + "Visualize: if I have three circles, each touching two others, it should form an equilateral triangle where the sides are of length 2r.\n", + "\n", + "Because each pair of circles, the distance between centers is the sum of radii, so 2r.\n", + "\n", + "So, the centers form an equilateral triangle with side 2r.\n", + "\n", + "Now, the shaded region: from the Asymptote code, it seems to be filling a area with a path that starts at (1,0) and goes down, etc., but it's a bit messy.\n", + "\n", + "\"dir(30)\" etc is Asymptote code for unit circle directions.\n", + "\n", + "Perhaps in the diagram, the shaded region is the set of points that are inside one of the circles, \n", + "Extracted:\n", + ",\n" + ] + }, + { + "data": { + "text/html": [ + "\n", + "
\n", + " \n", + " \n", + " [100/100 1:11:43, Epoch 0/1]\n", + "
\n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + "
StepTraining Lossrewardreward_stdcompletions / mean_lengthcompletions / min_lengthcompletions / max_lengthcompletions / clipped_ratiocompletions / mean_terminated_lengthcompletions / min_terminated_lengthcompletions / max_terminated_lengthklrewards / match_format_exactly / meanrewards / match_format_exactly / stdrewards / match_format_approximately / meanrewards / match_format_approximately / stdrewards / check_answer / meanrewards / check_answer / stdrewards / check_numbers / meanrewards / check_numbers / stdrewards / format_and_language_reward_func / meanrewards / format_and_language_reward_func / std
10.000000-1.5000004.618802843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0000000.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000001.0000004.618802
20.0000000.5000004.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0000000.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000003.0000004.000000
30.0000002.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0003180.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000005.0000000.000000
40.000000-5.8750000.750000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0009940.0000000.000000-0.5000000.000000-2.0000000.000000-0.3750000.750000-3.0000000.000000
50.0000001.7500000.866025843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0004780.0000000.000000-0.5000000.000000-2.0000000.000000-0.7500000.8660255.0000000.000000
60.000000-7.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0011320.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.000000-3.0000000.000000
70.000000-7.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0008670.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.000000-3.0000000.000000
80.0000001.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005780.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.0000005.0000000.000000
90.000000-7.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005500.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.000000-3.0000000.000000
100.000000-0.2500004.555217843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0007360.0000000.000000-0.5000000.000000-2.0000000.000000-0.7500000.8660253.0000004.000000
110.000000-7.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0006240.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.000000-3.0000000.000000
120.0000002.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005600.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000005.0000000.000000
130.000000-7.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005130.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.000000-3.0000000.000000
140.000000-5.0000004.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0008660.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.000000-1.0000004.000000
150.000000-5.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0010820.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.000000-3.0000000.000000
160.000000-4.6250004.750000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0006510.0000000.000000-0.5000000.000000-2.0000000.000000-1.1250000.750000-1.0000004.000000
170.000000-0.2500004.555217843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0008100.0000000.000000-0.5000000.000000-2.0000000.000000-0.7500000.8660253.0000004.000000
180.000000-1.8750004.230347843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0004390.0000000.000000-0.5000000.000000-2.0000000.000000-0.3750000.7500001.0000004.618802
190.000000-6.2500000.866025843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0009660.0000000.000000-0.5000000.000000-2.0000000.000000-0.7500000.866025-3.0000000.000000
200.0000002.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0007910.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000005.0000000.000000
210.000000-3.5000004.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0008160.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.000000-1.0000004.000000
220.000000-5.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0008440.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.000000-3.0000000.000000
230.000000-1.0000004.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0004250.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.0000003.0000004.000000
240.0000002.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0004540.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000005.0000000.000000
250.000000-0.2500003.570714843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0008240.0000000.000000-0.5000000.000000-2.0000000.000000-0.7500000.8660253.0000004.000000
260.000000-6.2500000.866025843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0004590.0000000.000000-0.5000000.000000-2.0000000.000000-0.7500000.866025-3.0000000.000000
270.000000-5.0000004.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0009520.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.000000-1.0000004.000000
280.000000-5.0000004.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005750.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.000000-1.0000004.000000
290.000000-5.8750000.750000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0018300.0000000.000000-0.5000000.000000-2.0000000.000000-0.3750000.750000-3.0000000.000000
300.000000-3.0000004.618802843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0006220.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.0000001.0000004.618802
310.0000001.7500000.866025843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0004440.0000000.000000-0.5000000.000000-2.0000000.000000-0.7500000.8660255.0000000.000000
320.000000-1.5000004.618802843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0007610.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000001.0000004.618802
330.000000-5.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0012550.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.000000-3.0000000.000000
340.000000-6.6250000.750000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0009640.0000000.000000-0.5000000.000000-2.0000000.000000-1.1250000.750000-3.0000000.000000
350.0000002.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0006730.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000005.0000000.000000
360.0000002.1250000.750000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0009360.0000000.000000-0.5000000.000000-2.0000000.000000-0.3750000.7500005.0000000.000000
370.000000-0.6250004.308422843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0004870.0000000.000000-0.5000000.000000-2.0000000.000000-1.1250000.7500003.0000004.000000
380.000000-7.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0012280.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.000000-3.0000000.000000
390.0000002.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0010430.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000005.0000000.000000
400.000000-7.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0004960.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.000000-3.0000000.000000
410.000000-6.2500000.866025843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005200.0000000.000000-0.5000000.000000-2.0000000.000000-0.7500000.866025-3.0000000.000000
420.0000000.5000004.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005280.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000003.0000004.000000
430.0000002.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005310.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000005.0000000.000000
440.000000-2.6250004.230347843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0006180.0000000.000000-0.5000000.000000-2.0000000.000000-1.1250000.7500001.0000004.618802
450.0000000.5000004.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0008380.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000003.0000004.000000
460.0000002.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0004650.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000005.0000000.000000
470.000000-7.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0010880.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.000000-3.0000000.000000
480.000000-3.5000004.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005590.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.000000-1.0000004.000000
490.0000002.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0004010.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000005.0000000.000000
500.000000-3.0000004.618802843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005260.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.0000001.0000004.618802
510.0000001.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005130.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.0000005.0000000.000000
520.000000-4.2500004.555217843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0007680.0000000.000000-0.5000000.000000-2.0000000.000000-0.7500000.866025-1.0000004.000000
530.000000-6.2500000.866025843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0010670.0000000.000000-0.5000000.000000-2.0000000.000000-0.7500000.866025-3.0000000.000000
540.0000001.7500000.866025843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0004720.0000000.000000-0.5000000.000000-2.0000000.000000-0.7500000.8660255.0000000.000000
550.0000000.5000004.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0009730.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000003.0000004.000000
560.000000-1.0000004.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0006300.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.0000003.0000004.000000
570.000000-7.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0010440.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.000000-3.0000000.000000
580.000000-2.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0006880.0000000.000000-0.5000000.000000-2.0000000.0000003.5000000.000000-3.0000000.000000
590.0000001.7500000.866025843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005520.0000000.000000-0.5000000.000000-2.0000000.000000-0.7500000.8660255.0000000.000000
600.0000001.3750000.750000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0004410.0000000.000000-0.5000000.000000-2.0000000.000000-1.1250000.7500005.0000000.000000
610.0000002.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005530.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000005.0000000.000000
620.0000002.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0007020.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000005.0000000.000000
630.0000001.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005100.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.0000005.0000000.000000
640.0000002.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0006190.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000005.0000000.000000
650.0000001.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0004050.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.0000005.0000000.000000
660.0000001.7500000.866025843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005130.0000000.000000-0.5000000.000000-2.0000000.000000-0.7500000.8660255.0000000.000000
670.0000001.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0004970.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.0000005.0000000.000000
680.000000-3.8750003.250000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0011870.0000000.000000-0.5000000.000000-2.0000000.000000-0.3750000.750000-1.0000004.000000
690.000000-5.0000004.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0006760.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.000000-1.0000004.000000
700.000000-7.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005800.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.000000-3.0000000.000000
710.0000000.1250004.750000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0008110.0000000.000000-0.5000000.000000-2.0000000.000000-0.3750000.7500003.0000004.000000
720.000000-3.5000004.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0007910.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.000000-1.0000004.000000
730.0000001.3750000.750000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005390.0000000.000000-0.5000000.000000-2.0000000.000000-1.1250000.7500005.0000000.000000
740.0000001.7500000.866025843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005320.0000000.000000-0.5000000.000000-2.0000000.000000-0.7500000.8660255.0000000.000000
750.000000-7.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0006650.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.000000-3.0000000.000000
760.0000001.7500004.173328843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0006740.0000000.000000-0.5000000.000000-2.0000000.0000001.7500002.0207262.5000005.000000
770.0000001.3750000.750000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0002960.0000000.000000-0.5000000.000000-2.0000000.000000-1.1250000.7500005.0000000.000000
780.000000-0.6250004.308422843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0006790.0000000.000000-0.5000000.000000-2.0000000.000000-1.1250000.7500003.0000004.000000
790.000000-5.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0010930.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.000000-3.0000000.000000
800.000000-0.6250003.250000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0009480.0000000.000000-0.5000000.000000-2.0000000.000000-1.1250000.7500003.0000004.000000
810.0000004.2500002.020726843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005160.0000000.000000-0.5000000.000000-2.0000000.0000001.7500002.0207265.0000000.000000
820.000000-5.0000004.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0006600.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.000000-1.0000004.000000
830.000000-6.6250000.750000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0007810.0000000.000000-0.5000000.000000-2.0000000.000000-1.1250000.750000-3.0000000.000000
840.000000-7.0000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0012930.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.000000-3.0000000.000000
850.0000002.1250000.750000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0005290.0000000.000000-0.5000000.000000-2.0000000.000000-0.3750000.7500005.0000000.000000
860.000000-5.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0007970.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.000000-3.0000000.000000
870.000000-5.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0011510.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.000000-3.0000000.000000
880.000000-3.0000004.618802843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0006250.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.0000001.0000004.618802
890.0000000.1250004.750000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0011700.0000000.000000-0.5000000.000000-2.0000000.000000-0.3750000.7500003.0000004.000000
900.0000001.7500000.866025843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0004320.0000000.000000-0.5000000.000000-2.0000000.000000-0.7500000.8660255.0000000.000000
910.000000-5.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0008810.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.000000-3.0000000.000000
920.000000-2.2500005.484828843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0011370.0000000.000000-0.5000000.000000-2.0000000.000000-0.7500000.8660251.0000004.618802
930.000000-0.6250003.250000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0007900.0000000.000000-0.5000000.000000-2.0000000.000000-1.1250000.7500003.0000004.000000
940.000000-3.5000004.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0008930.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.000000-1.0000004.000000
950.0000002.5000000.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0006660.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000005.0000000.000000
960.000000-2.2500005.484828843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0006860.0000000.000000-0.5000000.000000-2.0000000.000000-0.7500000.8660251.0000004.618802
970.000000-1.0000004.000000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0011330.0000000.000000-0.5000000.000000-2.0000000.000000-1.5000000.0000003.0000004.000000
980.000000-2.6250004.230347843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0006440.0000000.000000-0.5000000.000000-2.0000000.000000-1.1250000.7500001.0000004.618802
990.000000-1.5000004.618802843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0009740.0000000.000000-0.5000000.000000-2.0000000.0000000.0000000.0000001.0000004.618802
1000.0000002.1250000.750000843.000000843.000000843.0000001.0000000.0000000.0000000.0000000.0003650.0000000.000000-0.5000000.000000-2.0000000.000000-0.3750000.7500005.0000000.000000

" + ], + "text/plain": [ + "" + ] + }, + "metadata": {}, + "output_type": "display_data" + }, + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Unsloth: Will smartly offload gradients to save VRAM!\n", + "********************Question:\n", + "In a class of 20 students, all but 4 of the students put their names on a typed assignment. If the teacher randomly guesses, what is the probability that she correctly guesses which paper belongs to each of the four remaining students? Express your answer as a common fraction.The answer is in the form \f\n", + "rac{m}{n}, where gcd(m, n) = 1. Please provide the value of m + n. \n", + "Answer:\n", + "25 \n", + "Response:\n", + "\n", + "The problem states there are 20 students, and all but 4 put their names on their assignments. So, 16 students have their names on their papers, and 4 do not. The teacher guesses randomly, and I need to find the probability that she correctly identifies which paper belongs to each of the four students who didn't put their names on.\n", + "\n", + "First, I need to understand what's being asked. The teacher is guessing which paper is whose for these four students, but since the papers are unlabeled, I assume she's assigning labels to the papers or something. Let me read the problem again.\n", + "\n", + "\"In a class of 20 students, all but 4 of the students put their names on a typed assignment. If the teacher randomly guesses, what is the probability that she correctly guesses which paper belongs to each of the four remaining students?\"\n", + "\n", + "So, there are 20 papers, all identical except for the names, but four don't have names. The teacher needs to guess the identity of these four papers, meaning which paper corresponds to which student who didn't name it.\n", + "\n", + "But the papers are indistinguishable except for the ones with names? Or are they distinct? The papers are probably distinct because they are from different students, even though some don't have names. So, there are 20 distinct papers, 16 with names labeled, and 4 without.\n", + "\n", + "The teacher is guessing the assignment for these four unlabeled papers to the four students who didn't label them.\n", + "\n", + "But since the four unlabeled papers are indistinguishable in terms of the teacher's guessing, I need to think about how she is guessing.\n", + "\n", + "Probably, the teacher is considering all papers and trying to assign each to a student, but since there are students who did label their papers, and others didn't.\n", + "\n", + "This is a bit confusing. Let me model it.\n", + "\n", + "There are 20 distinct papers.\n", + "\n", + "16 of them have the correct student's name on them.\n", + "\n", + "4 of them do not have any name.\n", + "\n", + "The teacher needs to guess which paper is whose for all students, but the question specifically asks for the probability that she correctly guesses which paper belongs to each of the four remaining students. So, focus is on the unlabeled papers.\n", + "\n", + "But for the unlabeled papers, they need to be matched to the four students who didn't label their own.\n", + "\n", + "But the teacher is guessing randomly, I assume by assigning students to papers or papers to students.\n", + "\n", + "Perhaps it's better to think that the teacher is trying to permute the assignments, but with the labeled papers fixed or something.\n", + "\n", + "Another way to think: the teacher is making a random guess for the entire assignment, but since some papers are already labeled, maybe those are fixed.\n", + "\n", + "The problem doesn't specify that the labeled papers are correctly identified or not. Let's read carefully.\n", + "\n", + "The teacher is guessing the entire assignment, but the condition is that she guesses correctly only for the four unlabeled papers, and for the labeled ones, it might be correct or not, but the question is only about the four unlabeled students.\n", + "\n", + "The probability is for her to correctly guess which paper belongs to each of the four remaining students, meaning that for those four specific students, the paper assigned by her guess matches the actual paper.\n", + "\n", + "But when she guesses, she might be guessing the entire permutation.\n", + "\n", + "Since the labeled papers are already named, perhaps when she guesses, she is guessing which paper is which student, including the unlabeled ones.\n", + "\n", + "But the labeled papers might or might not be correctly labeled; the problem doesn't specify.\n", + "\n", + "To simplify, perhaps we can think that the teacher is assigning students to papers randomly.\n", + "\n", + "But for the papers that are already labeled, her guess must match the label, but the label is correct by assumption, since the student put their own name.\n", + "\n", + "Let's clarify.\n", + "\n", + "Each paper has a fixed correct student.\n", + "\n", + "16 papers have the correct student's name written on them.\n", + "\n", + "4 papers do not have any name written.\n", + "\n", + "The teacher tries to guess, by assigning each paper to a student, i.e., she assigns a random student to each paper, or equivalently, a random permutation.\n", + "\n", + "Then, for the four students \n", + "Extracted:\n", + "20\n", + "********************Question:\n", + "Find the minimum possible value of\n", + "\\[\\sqrt{58 - 42x} + \\sqrt{149 - 140 \\sqrt{1 - x^2}}\\]where $-1 \\le x \\le 1.$The answer is in the form k\\sqrt{m}+n,. Please provide the value of k + m + n. \n", + "Answer:\n", + "110 \n", + "Response:\n", + "\n", + "I need to find the minimum value of \\(\\sqrt{58 - 42x} + \\sqrt{149 - 140 \\sqrt{1 - x^2}}\\) for \\(x\\) in \\([-1, 1]\\), and the answer is in the form \\(k\\sqrt{m} + n\\), and then find \\(k + m + n\\).\n", + "\n", + "First, I notice that there's a \\(\\sqrt{1 - x^2}\\) term, which suggests a trigonometric substitution might help. Let me set \\(x = \\cos \\theta\\), so \\(\\sqrt{1 - x^2} = \\sqrt{1 - \\cos^2 \\theta} = \\sin \\theta\\), assuming \\(\\theta\\) is in \\([0, \\pi]\\) since \\(x \\in [-1, 1]\\).\n", + "\n", + "If \\(x = \\cos \\theta\\), then \\(\\sqrt{1 - x^2} = \\sin \\theta\\) for \\(\\theta \\in [0, \\pi]\\).\n", + "\n", + "Now, the expression becomes:\n", + "\n", + "\\(\\sqrt{58 - 42 \\cos \\theta} + \\sqrt{149 - 140 \\sin \\theta}\\)\n", + "\n", + "Because \\(\\sqrt{1 - x^2} = \\sin \\theta\\), so \\(149 - 140 \\sqrt{1 - x^2} = 149 - 140 \\sin \\theta\\).\n", + "\n", + "Now, I need to minimize this sum with respect to \\(\\theta\\).\n", + "\n", + "Sums of square roots often represent distances, so maybe this can be interpreted geometrically.\n", + "\n", + "Consider \\(\\sqrt{a - b \\cos \\theta} + \\sqrt{c - d \\sin \\theta}\\). This might be related to distances in the plane.\n", + "\n", + "I recall that \\(\\sqrt{a - b \\cos \\theta}\\) can be thought of as a distance involving angles.\n", + "\n", + "Set \\(A = \\sqrt{58 - 42x}\\), and \\(B = \\sqrt{149 - 140 \\sqrt{1 - x^2}}\\).\n", + "\n", + "But \\(B\\) has a nested square root, which might be messy.\n", + "\n", + "With trigonometric substitution, so \\(x = \\cos \\theta\\), expression is \\(\\sqrt{58 - 42 \\cos \\theta} + \\sqrt{149 - 140 \\sin \\theta}\\).\n", + "\n", + "Let me simplify the expressions inside the square roots.\n", + "\n", + "First, \\(\\sqrt{58 - 42 \\cos \\theta}\\). I can write this as \\(\\sqrt{a} \\sqrt{b - c \\cos \\theta}\\), but perhaps better to factor out constants.\n", + "\n", + "Notice that 58 and 42 might have a common factor. 58 and 42, gcd is 2. So 42 = 2*21, 58 = 2*29, so \\(\\sqrt{2(29 - 21 \\cos \\theta)}\\).\n", + "\n", + "Similarly for the other term, 149 and 140, gcd? 149 is prime? 149 divided by 11 is about 13.5, 13*11=143, 149-143=6, not divisible. 149 ÷ 149=1, so probably prime. 140=2^3*5*7, no common factors, so \\(\\sqrt{149 - 140 \\sin \\theta}\\).\n", + "\n", + "But in the first term, after factoring out 2, it's \\(\\sqrt{2} \\sqrt{29 - 21 \\cos \\theta}\\).\n", + "\n", + "Now, 29 - 21 cos θ. I can write this as R - S cos θ.\n", + "\n", + "Generally, expressions like \\(\\sqrt{a - b \n", + "Extracted:\n", + "58\n", + "********************Question:\n", + "Find the number of three-element subsets of $\\{1, 2, 3,...,13\\}$ that contain at least one element that is a multiple of $2$, at least one element that is a multiple of $3$, and at least one element that is a multiple of $5$ such as $\\{2,3, 5\\}$ or $\\{6, 10,13\\}$. \n", + "Answer:\n", + "63 \n", + "Response:\n", + "\n", + "The problem is to find the number of three-element subsets of {1, 2, 3, ..., 13} that contain at least one multiple of 2, at least one multiple of 3, and at least one multiple of 5. Examples given are {2,3,5} and {6,10,13}.\n", + "\n", + "First, I need to find all subsets of size three that satisfy these three conditions. Since the conditions are about the presence of multiples, I should think in terms of inclusion or maybe complementary counting.\n", + "\n", + "Let me list out the multiples of 2, 3, and 5 within 1 to 13.\n", + "\n", + "Multiples of 2: 2,4,6,8,10,12 → there are 6 of them.\n", + "\n", + "Multiples of 3: 3,6,9,12,15? 15 is greater than 13, so 3,6,9,12 → but 12 is included, yes. So {3,6,9,12} → 4 numbers.\n", + "\n", + "Multiples of 5: 5,10,15? 15>13, so {5,10} → 2 numbers.\n", + "\n", + "But these sets overlap because some numbers are multiples of more than one.\n", + "\n", + "The condition is that in the subset, there must be at least one multiple of 2, one of 3, and one of 5.\n", + "\n", + "Since the subset has three elements, and the conditions must be satisfied, I need to ensure that not all conditions can be missed.\n", + "\n", + "This sounds like a job for the inclusion-exclusion principle or perhaps complementary counting.\n", + "\n", + "Let me denote A as the set of elements that are multiples of 2.\n", + "\n", + "B for multiples of 3.\n", + "\n", + "C for multiples of 5.\n", + "\n", + "But in subset selection, it's about the subset containing at least one from each.\n", + "\n", + "I could calculate the total number of three-element subsets, and subtract those that miss at least one of the conditions.\n", + "\n", + "Total three-element subsets of {1,2,...,13} is C(13,3).\n", + "\n", + "C(13,3) = 13*12*11 / 6 = 286? Let me compute: 13*12=156, 156/2=78, no: C(13,3) = 13! / (3! * 10!) = (13*12*11)/(3*2*1) = (13*12*11)/6.\n", + "\n", + "13*12=156, 156/6=26? 156 divided by 6 is 26, then 26*11=286, yes.\n", + "\n", + "So total subsets: 286.\n", + "\n", + "Now, I need to subtract the subsets that do not have at least one multiple of 2, or no multiple of 3, or no multiple of 5.\n", + "\n", + "Using complementary counting: the subsets that fail at least one condition.\n", + "\n", + "So, by inclusion-exclusion, the number of subsets that miss at least one condition is equal to the sum over each singleton condition of the subsets missing that condition, minus the sum over intersections (missing two conditions), plus the subsets that miss all three conditions.\n", + "\n", + "The condition is that the subset must contain at least one from A, one from B, one from C.\n", + "\n", + "So, the subsets that do not satisfy this are those that:\n", + "\n", + "- Lack at least one multiple of 2, or\n", + "\n", + "- Lack at least one multiple of 3, or\n", + "\n", + "- Lack at least one multiple of 5.\n", + "\n", + "But \"lack at least one\" means that for at least one prime (2,3,5), there is no multiple in the subset.\n", + "\n", + "For the condition not to hold, it means that for at least one of the three properties, the subset does not have any element satisfying that property.\n", + "\n", + "So, to use inclusion-exclusion, I need to \n", + "Extracted:\n", + "1,\n", + "********************Question:\n", + "The graph of the equation $x^{y}=y^{x}$ in the first quadrant (i.e., the region where $x>0$ and $y>0$) consists of a straight line and a curve. Find the sum of the coordinates of the intersection points of a straight line and a curve, rounded down. \n", + "Answer:\n", + "5 \n", + "Response:\n", + "\n", + "Wah, pertama kali menemukan soal ini cukup menantang. Persamaannya adalah \\(x^y = y^x\\), dan kita diminta untuk mencari titik potong antara grafiknya dengan sumbu-sumbunya atau apakah itu dengan sumbu lain? Sepertinya ada kesalahpahalan dalam soal karena yang diminta adalah \"intersection points of a straight line and a curve\", tapi tanpa spesifikasinya garis yang mana.\n", + "\n", + "Mungkin yang dimaksud adalah titik potong grafik dengan sumbu-sumbu koordinat. Dalam konteks pertama kuadrant, artinya x dan y positif. Persamaan \\(x^y = y^x\\) ini sebenarnya bisa diselesaikan untuk mencari titik potong pada sumbu x dan y.\n", + "\n", + "Pertama, mari kita pikirkan grafiknya. Persamaan ini simetris, artinya jika (a,b) adalah solusi, maka (b,a) juga solusi. Dengan a dan b positif.\n", + "\n", + "Untuk menemukan mana yang merupakan garis dan mana yang merupakan kurva. Ya, ada garis lurus y=x, karena jika x=y, maka persamaannya menjadi identik, yaitu x^x = x^x, jadi selalu benar untuk semua x>0. Jadi, tadi titik potong antara x=y dengan kurva lainnya.\n", + "\n", + "Tunggu, tidak. Persamaan \\(x^y = y^x\\) sendiri secara aljabar tidak langsung menghasilkan grafik garis lurus. Mungkin titik potongnya adalah ketika salah satu variabel adalah 1 atau beberapa nilai khusus.\n", + "\n", + "Kita bisa menyelesaikan persamaan ini untuk mencari solusi. Misalnya, jika x=1, maka \\(1^y = y^1\\), jadi 1 = y, jadi y=1. Jadi titik (1,1) adalah solusinya.\n", + "\n", + "Jika y=1, sama saja, x bisa sembarang? Tidak, karena jika y=1, maka \\(x^1 = 1^x\\), jadi x =1. Jadi titik (1,1) adalah satu-satunya titik untuk garis x=1 dan y=1? Tidak, garis x=1 adalah solusi hanya ketika y=1, tapi tidak, garis x=1, y bisa apa saja? Tidak, karena persamaan x^y = y^x.\n", + "\n", + "Mari kita sesuaikan. Misalnya, jika x=1, maka 1^y = y^x, tapi 1^y =1 untuk semua y, dan y^x = y^1 = y, jadi 1 = y, jadi hanya (1,1).\n", + "\n", + "Jika y=1, maka x^1 = 1^x, jadi x =1, sama.\n", + "\n", + "Jika x=y, maka x^x = x^x, benar untuk semua x>0. Tapi ini bukan garis lurus di grafik, karena grafik x^y = y^x tidak identik dengan x=y kecuali di titik-titik tertentu.\n", + "\n", + "Ketika x=y, benar, tapi ada lebih banyak solusi. Misalnya, x=4, y=2: 4^2 =16, 2^4=16, benar. Jadi (4,2) adalah solusi. Dan juga (2,4) karena simetris.\n", + "\n", + "Jadi ada beberapa titik potong. Tapi di mana garis lurus dan kurva itu? Mungkin salah satu bagian grafik ini lurus dan bagian lain tidak.\n", + "\n", + "Mari kita ganti variabel. Seringkali, untuk memecahkan x^y = y^x, kita bisa gunakan logar \n", + "Extracted:\n", + ",\n", + "********************Question:\n", + "Triangle $\\triangle ABC$ has $AB= 3$, $BC = 4$, and $AC = 5$. Let $M$ and $N$ be the midpoints of $AC$ and $BC$, respectively. If line $AN$ intersects the circumcircle of triangle $\\triangle BMC$ at points $X$ and $Y$, then $XY^2 = \\frac{m}{n}$ for some relatively prime positive integers $m,n$. Find $m+n$. \n", + "Answer:\n", + "933 \n", + "Response:\n", + "\n", + "The problem involves a triangle with sides AB=3, BC=4, AC=5. I notice that 3-4-5 is a Pythagorean triple, so triangle ABC is a right-angled triangle with right angle at B, since 3² + 4² = 9 + 16 = 25 = 5². So angle B is 90 degrees.\n", + "\n", + "M and N are midpoints of AC and BC respectively. Since it's a right triangle with right angle at B, I should probably draw a diagram.\n", + "\n", + "Sketch triangle ABC with AB=3, BC=4, AC=5, right-angled at B. So, put B at origin (0,0), C at (4,0) since BC=4, and A at (0,3) since AB=3. Distance from A to C: from (0,3) to (4,0), distance is sqrt((4-0)^2 + (0-3)^2) = sqrt(16+9)=5, yes.\n", + "\n", + "Now, M is midpoint of AC. A(0,3), C(4,0), so midpoint M: ((0+4)/2, (3+0)/2) = (2, 1.5)\n", + "\n", + "1.5 is 3/2, so M(2, 3/2)\n", + "\n", + "N is midpoint of BC. B(0,0), C(4,0), so N(2, 0)\n", + "\n", + "Now, line AN: A is (0,3), N is (2,0)\n", + "\n", + "So, let's find the equation of line AN.\n", + "\n", + "Points A(0,3), N(2,0)\n", + "\n", + "Slope m_AN = (0 - 3)/(2 - 0) = -3/2\n", + "\n", + "So slope is -3/2\n", + "\n", + "Equation: y - 3 = m(x - 0) => y = - (3/2)x + 3\n", + "\n", + "Now, we need to find the circumcircle of triangle BMC.\n", + "\n", + "First, find points B, M, C.\n", + "\n", + "B(0,0), M(2,3/2), C(4,0)\n", + "\n", + "So, triangle with vertices (0,0), (2,1.5), (4,0)\n", + "\n", + "Now, circumcircle of triangle BMC.\n", + "\n", + "Since it's a triangle, we need to find the circle passing through these three points.\n", + "\n", + "Then, find the intersection of line AN with this circumcircle, and find the length XY, but it says intersects at points X and Y, and find XY², and it should be m/n fraction.\n", + "\n", + "Since it's a circle and a line intersecting at two points, I need to find the chord length or something.\n", + "\n", + "I recall that for two points on a circle, the length of the chord can be found using the circle equation and the line.\n", + "\n", + "But maybe there's a better way using geometry properties.\n", + "\n", + "Since it's a circle and a line, I can find the equation of the circle and the line, find intersection points, then distance between them.\n", + "\n", + "But with coordinates, I can do that.\n", + "\n", + "First, let's find the circumcircle of B, M, C.\n", + "\n", + "Points: B(0,0), C(4,0), M(2, 3/2)\n", + "\n", + "So, three points: (0,0), (4,0), (2, 1.5)\n", + "\n", + "To find circumcircle, solve for circle equation.\n", + "\n", + "General circle equation: x² + y² + Dx + Ey + F = 0\n", + "\n", + "Plug in points.\n", + "\n", + "First, B(0,0): 0 + 0 + 0 + 0 + F = 0 => F = 0\n", + "\n", + "F=0, since passes through (0,0)\n", + "\n", + "Now, C(4,0): 16 + 0 + 4D + 0 + F = 0\n", + "\n", + "D and F, but F=0, so 16 + 4D = 0 => 4D = -16 => D = -4\n", + "\n", + " \n", + "Extracted:\n", + "3,\n", + "********************Question:\n", + "Find the product of uncommon real roots of the two polynomials $x^4 + 2x^3 - 8x^2 - 6x + 15$ and $x^3 + 4x^2 - x - 10$. \n", + "Answer:\n", + "6 \n", + "Response:\n", + "\n", + "Soalnya adalah mencari hasil kali akar tidak umum dari dua polinomial: \\( x^4 + 2x^3 - 8x^2 - 6x + 15 \\) dan \\( x^3 + 4x^2 - x - 10 \\). Saya perlu mencari akar-akar yang sama, tapi karena ini adalah polinomial, mungkin lebih baik mencari faktor persekutuan.\n", + "\n", + "Pertama-tama, saya akan memfaktorkan kedua polinomial untuk mencari akar-akar dan kemudian mencari akar yang sama.\n", + "\n", + "Mulai dengan polinomial kedua yang lebih sederhana: \\( x^3 + 4x^2 - x - 10 \\).\n", + "\n", + "Saya harus mencari akar-akarnya. Karena ini polinomial kubik, akar rasionalnya mungkin berupa faktor dari konstanta dibagi dengan faktor dari koefisien utama. Konstanta adalah -10, koefisien utama 1, jadi kemungkinan akar rasionalnya adalah ±1, ±2, ±5, ±10.\n", + "\n", + "Coba \\( x = 1 \\): \\( 1 + 4 - 1 - 10 = -6 \\), bukan 0.\n", + "\n", + "\\( x = -1 \\): \\( -1 + 4 + 1 - 10 = -6 \\), bukan 0.\n", + "\n", + "\\( x = 2 \\): \\( 8 + 4*4 - 2 - 10 = 8 + 16 - 2 - 10 = 12 \\), bukan 0.\n", + "\n", + "\\( x = -2 \\): \\( -8 + 4*4 - (-2) - 10 = -8 + 16 + 2 - 10 = 0 \\). Ya! Jadi \\( x = -2 \\) adalah akar.\n", + "\n", + "Jadi, \\( x + 2 \\) adalah faktornya. Saya akan melakukan pembagian polinomial.\n", + "\n", + "Bagi \\( x^3 + 4x^2 - x - 10 \\) dengan \\( x + 2 \\).\n", + "\n", + "Pembagian polinomial:\n", + "\n", + "Pertama, \\( x^3 \\) dibagi \\( x \\) = \\( x^2 \\), lalu \\( x^2 \\) dikalikan \\( x + 2 \\) = \\( x^3 + 2x^2 \\), kurangkan dari polinomial: \\( (x^3 + 4x^2 - x - 10) - (x^3 + 2x^2) = 2x^2 - x - 10 \\).\n", + "\n", + "Saya pikir saya salah. Polinomial asli adalah \\( x^3 + 4x^2 - x - 10 \\), dikurangi \\( x^3 + 2x^2 \\), jadi \\( (x^3 + 4x^2) - (x^3 + 2x^2) = 2x^2 \\), kemudian -x, -10, jadi hasil sementara \\( 2x^2 - x - 10 \\).\n", + "\n", + "Sekarang, \\( 2x^2 - x - 10 \\) dibagi \\( x + 2 \\). \\( 2x^2 / x = 2x \\), \\( 2x * (x + 2) = 2x^2 + 4x \\), kurangkan: \\( (2x^2 - x) - (2x^2 + 4x) = -5x \\), lalu -10.\n", + "\n", + "Jadi \\( -5x - 10 \\), dibagi \\( x + 2 \\), \\( -5 \\), \\( -5 * (x + 2) = -5 \n", + "Extracted:\n", + "4\n", + "********************Question:\n", + "A Parliament of $2000 MP$s decided to ratify the state budget of 200 expenditure items. Each $MP$ prepared a draft budget with what (s)he thinks the maximum possible allocation for each item so that the total expenditure does not exceed a given ceiling, S. For each item, the Parliament approves the maximum expenditure approved by not less than k $MP$s. What is the least value of k to ensure that the approved total does not exceed S? \n", + "Answer:\n", + "1991 \n", + "Response:\n", + "\n", + "Pertama, permasalahan ini tentang anggota parlemen (MP) yang membuat draf budget untuk 2000 item pengeluaran. Setiap MP memberikan alokasi maksimum untuk setiap item sehingga total pengeluaran tidak melebihi S.\n", + "\n", + "Tujuannya adalah menemukan nilai k minimum sehingga ketika parlemen menyetujui alokasi maksimum untuk suatu item hanya jika setidaknya k MP menyetujui alokasi tersebut, maka total pengeluaran yang disetujui tidak akan melebihi S.\n", + "\n", + "Kata lain, kita perlu memastikan bahwa setiap item yang disetujui memiliki alokasi yang tidak terlalu tinggi sehingga ketika dijumlahkan, totalnya tidak melebihi S, dengan asumsi bahwa masing-masing MP mengerahkan alokasi maksimum yang mungkin tanpa melampaui S.\n", + "\n", + "Setiap MP membuat draf budget dengan alokasi untuk masing-masing item sehingga total tidak melebihi S. Jadi, untuk setiap MP, jumlah alokasi item-itemnya adalah ≤ S.\n", + "\n", + "Kemudian, untuk setiap item j, alokasi yang disetujui adalah maksimum di antara alokasi yang diajukan oleh setidaknya k MP.\n", + "\n", + "Tapi yang disetujui adalah alokasi tertinggi yang diajukan oleh k atau lebih banyak MP.\n", + "\n", + "Jadi, untuk item j, kita cari alokasi tertinggi yang muncul pada k atau lebih banyak draf.\n", + "\n", + "Kemudian, kita jumlahkan semua alokasi-item untuk item-item tersebut.\n", + "\n", + "Dan kita ingin memastikan bahwa totalnya ≤ S.\n", + "\n", + "Tapi wait, ketika kita menyetujui alokasi untuk setiap item, alokasi yang disetujui adalah yang tertinggi dari MP yang menyetujudengan alokasi tersebut.\n", + "\n", + "Tapi dari perspektiva alokasi, kita perlu memastikan bahwa meskipun beberapa MP mungkin memiliki alokasi tinggi untuk satu item, tetapi dengan k cukup tinggi, kita hanya mengambil alokasi yang diajukan oleh banyak MP, jadi alokasi yang disetujui lebih rendah.\n", + "\n", + "Tapi tujuannya adalah menemukan k minimal sehingga setuju dengan skema ini, total alokasi yang disetujui tidak melebihi S.\n", + "\n", + "Sekarang, pertanyaannya adalah: dari situ, bagaimana menemukan k minimal agar pasti total tidak melebihi S.\n", + "\n", + "Karena setiap MP, draf dia memenuhi total ≤ S, tetapi dia bisa mengajukan alokasi tinggi untuk satu item dan rendah untuk lainya.\n", + "\n", + "Kita ingin mencocokkan seberapa tinggi alokasi tertinggi untuk setiap item saat disetujui.\n", + "\n", + "Mari kita misalkan untuk setiap item j, ada alokasi tertinggi α_j yang diajukan oleh s_j MP.\n", + "\n", + "Tapi yang disetujui adalah maks α_j yang diajukan oleh ≥ k MP.\n", + "\n", + "Atau, lebih tepatnya, alokasi yang disetujui untuk item j adalah supremum over α_j dengan s_j ≥ k.\n", + "\n", + "Ya.\n", + "\n", + "Sekarang, kita perlu memastikan bahwa jumlah dari semua α_tersetujui ≤ S.\n", + "\n", + "Tetapi ini tidak selalu benar untuk k kecil; jika k kecil, mungkin ada item yang banyak MP setuju dengan alokasi tinggi, jadi total bisa melebihi S.\n", + "\n", + "Jika k besar, maka hanya alokasi yang diajukan oleh banyak MP yang akan disetujui, jadi alokasi terendah dari kelompok besar, dan total mungkin lebih rendah.\n", + "\n", + "Tetapi S adalah batas atas, jadi jika kita mengurangi alokasi yang disetujui, \n", + "Extracted:\n", + ",\n", + "********************Question:\n", + "Let $T$ be the set of all positive integer divisors of $2004^{100}$. What is the largest possible number of elements of a subset $S$ of $T$ such that no element in $S$ divides any other element in $S$? \n", + "Answer:\n", + "10201 \n", + "Response:\n", + "\n", + "The problem states: Let T be the set of all positive integer divisors of 2004^100. I need to find the largest possible number of elements in a subset S of T such that no element in S divides any other element in S. I must think in Bahasa Indonesia, but that doesn't affect my solution.\n", + "\n", + "First, I need to understand what T is. T is the set of all divisors of 2004^100. To find the divisors, I should start by factorizing 2004.\n", + "\n", + "2004 divided by 2 is 1002\n", + "\n", + "1002 divided by 2 is 501\n", + "\n", + "501 divided by 3 is 167\n", + "\n", + "167 is a prime number. Let me confirm: 167, is it prime? It's not divisible by 2,3,5,7,11,13. 13*12=156, 167-156=11, not divisible. So yes, prime.\n", + "\n", + "So, 2004 = 2^2 * 3^1 * 167^1\n", + "\n", + "Therefore, 2004^100 = (2^2 * 3^1 * 167^1)^100 = 2^{200} * 3^{100} * 167^{100}\n", + "\n", + "Now, the number of divisors of a number is found by adding one to each of the exponents in the prime factorization and then multiplying.\n", + "\n", + "So, for 2004^100, the exponents are 200 for 2, 100 for 3, and 100 for 167.\n", + "\n", + "So, total number of divisors is (200+1)(100+1)(100+1) = 201 * 101 * 101\n", + "\n", + "But I don't need the total number yet. I need to find a subset where no one divides another.\n", + "\n", + "This sounds familiar; it's like finding an antichain in the divisor lattice.\n", + "\n", + "In lattice theory, an antichain is a set where no element is less than or equal to another in the partial order, which in this case is the divisibility order.\n", + "\n", + "And by Sperner's theorem or something, but I think for divisor lattices, the size of the largest antichain is given by the largest rank in the lattice, but I need to recall.\n", + "\n", + "I remember that for the divisor lattice of n = p1^{a1} p2^{a2} ... pk^{ak}, the largest antichain is the set of divisors with the median size, i.e., with the product of the primes raised to the floor or something average.\n", + "\n", + "In this case, the divisors can be thought of as vectors in the exponents space.\n", + "\n", + "Each divisor is determined by its exponents for each prime.\n", + "\n", + "So, for 2^{a} * 3^{b} * 167^{c}, where 0 ≤ a ≤ 200, 0 ≤ b ≤ 100, 0 ≤ c ≤ 100.\n", + "\n", + "The total number of divisors is 201 * 101 * 101, as I said.\n", + "\n", + "Now, the divisibility order: d1 divides d2 if for each prime, the exponent in d1 is less than or equal to that in d2.\n", + "\n", + "To have no two elements where one divides the other, it means that no two elements are comparable in this partial order.\n", + "\n", + "To maximize the size of an antichain.\n", + "\n", + "In the boolean lattice, Sperner's theorem applies, but here the exponents are not all the same size.\n", + "\n", + "Each exponent has different ranges.\n", + "\n", + "I recall that in such a product lattice, the size of the largest antichain is given by the maximum number of \n", + "Extracted:\n", + "2004\n", + "********************Question:\n", + "Evaluate the expression \\[(751 - 745) + (748 - 742) + (745 - 739) + (742 - 736) + \\cdots + (499 - 493) + (496 - 490).\\] \n", + "Answer:\n", + "516 \n", + "Response:\n", + "\n", + "Pertama-tama, saya harus memahami pola dalam ekspresi ini. Ekspresi ini terdiri dari serangkaian pasangan pengurangan, seperti (751 - 745), (748 - 742), dan seterusnya, hingga (496 - 490). Saya perlu mengevaluasi total dari semuanya itu.\n", + "\n", + "Mari kita lihat angka-angkanya. Pada awalnya ada 751 dan 745, lalu 748 dan 742, kemudian 745 dan 739, dan seterusnya. Tampaknya ada pola tertentu di sini.\n", + "\n", + "Coba lihat selisih dalam setiap pasangan:\n", + "\n", + "Untuk (751 - 745) = 6\n", + "\n", + "(748 - 742) = 6\n", + "\n", + "(745 - 739) = 6\n", + "\n", + "(742 - 736) = 6\n", + "\n", + "Semuanya memberikan selisih 6. Jadi, setiap pasangan menghasilkan 6. Tapi ini hanya untuk awalnya. Bagaimana dengan pasangan terakhir?\n", + "\n", + "Ekspresi tersebut berakhir dengan (496 - 490) = 6. Jadi, selalu 6?\n", + "\n", + "Tapi apakah ini berlaku untuk semua pasangan? Mari kita lihat pola angkanya.\n", + "\n", + "Perhatikan bahwa angka pertama dalam setiap pasangan adalah: 751, 748, 745, 742, ... dan angka kedua: 745, 742, 739, 736, ...\n", + "\n", + "Mari kita cari pola dari angka pertama dan angka kedua.\n", + "\n", + "Setiap angka pertama berkurang 3 setiap kali: 751, 748, 745, 742, ... Jadi, deret aritmatika dengan suku pertama a = 751, dan beda b = -3.\n", + "\n", + "Setiap angka kedua juga berkurang 3: 745, 742, 739, 736, ... Suku pertama c = 745, beda b = -3.\n", + "\n", + "Tapi, inisialnya berbeda. Pada (751 - 745), angka pertama 751, angka kedua 745. Lalu, (748 - 742), angka pertama 748 (yang 3 lebih kecil dari 751), angka kedua 742 (yang juga 3 lebih kecil dari 745).\n", + "\n", + "Untuk setiap pasangan, angka pertama dan angka kedua menurun sebesar 3.\n", + "\n", + "Jadi, untuk setiap n, pasangan (a_n - b_n), tapi tidak jelas dari mana.\n", + "\n", + "Mari kita pikirkan secara berbeda. Ekspresi ini mungkin dapat ditulis sebagai jumlah dari selisih antara angka-angka tertentu.\n", + "\n", + "Alethia mengamati bahwa setiap pasangan mengurangkan angka yang lebih kecil dari angka sebelumnya.\n", + "\n", + "Perhatikan bahwa dalam setiap pasangan seperti (751 - 745), 745 adalah hasil dari (751 - 6)? Tidak, 751 - 6 = 745, benar.\n", + "\n", + "751 - 745 = 6, dan 745 = 751 - 6.\n", + "\n", + "Tapi untuk pasangan kedua: angka pertama adalah 748, angka kedua 74 \n", + "Extracted:\n", + ",\n", + "********************Question:\n", + "Find the eighth term of the sequence $1440,$ $1716,$ $1848,\\ldots,$ whose terms are formed by multiplying the corresponding terms of two arithmetic sequences.\n", + " \n", + "Answer:\n", + "348 \n", + "Response:\n", + "\n", + "Saya harus menemukan suku kedelapan dari sebuah urutan yang diberikan: 1440, 1716, 1848, dan seterusnya. Urutannya terbentuk dengan mengalikan suku-sukunya dari dua barisan aritmatika. Saya harus berpikir dalam Bahasa Indonesia, tetapi saya akan menyelesaikannya dalam bahasa Inggris karena itu lebih mudah untuk saya.\n", + "\n", + "Pertama, saya perlu memahami apa itu barisan aritmatika. Barisan aritmatika adalah barisan angka yang memiliki beda yang sama antar suku-sukunya. Misalnya, 2, 5, 8, 11, ... adalah barisan aritmatika dengan beda 3.\n", + "\n", + "Dalam masalah ini, ada dua barisan aritmatika, dan setiap suku dari barisan hasil adalah hasil kali satu suku dari barisan pertama dan satu suku dari barisan kedua. Jadi, jika saya punya barisan A: a, a+d, a+2d, a+3d, ... dan barisan B: b, b+d, b+2d, b+3d, ... maka barisan hasilnya adalah ab, a(b+d), (a+d)(b+d), (a+d)(b+2d), dan seterusnya.\n", + "\n", + "Saya diberikan tiga suku pertama dari barisan hasil: 1440, 1716, 1848. Saya perlu mencari suku kedelapan.\n", + "\n", + "Saya akan mendasarkan diri pada tiga suku yang diberikan.\n", + "\n", + "Misalkan barisan pertama adalah A_n = a + (n-1)d\n", + "\n", + "Barisan kedua adalah B_n = b + (n-1)e\n", + "\n", + "Kemudian suku ke-n dari barisan hasil adalah P_n = A_n * B_n = [a + (n-1)d] * [b + (n-1)e]\n", + "\n", + "Sekarang, untuk n=1: P1 = (a)(b) = 1440\n", + "\n", + "Untuk n=2: P2 = (a+d)(b+e) = 1716\n", + "\n", + "Untuk n=3: P3 = (a+2d)(b+2e) = 1848\n", + "\n", + "Saya memiliki tiga persamaan tetapi empat variabel: a, b, d, e. Itu tidak akan cukup dengan hanya tiga persamaan. Mungkin d dan e adalah beda, dan saya perlu mencari hubungan.\n", + "\n", + "Saya juga perlu mencari P8, yang berarti saya perlu mencari nilai dari [a + 7d] * [b + 7e]\n", + "\n", + "Tapi dengan empat variabel dan hanya tiga persamaan, saya mungkin perlu beberapa asumsi atau cara lain.\n", + "\n", + "Cara lain: mungkin kedua barisan aritmatika memiliki beda yang sama, atau sesuatu. Tapi masalah tidak mengatakan itu; itu hanya mengatakan dua barisan aritmatika.\n", + "\n", + "Jika beda sama, misalnya d dan e sama, maka bisa lebih mudah. Saya bisa mencoba asumsikan itu.\n", + "\n", + "Jika d = e, maka barisan hasilnya akan berbeda.\n", + "\n", + "Misalkan d = e.\n", + "\n", + "Maka untuk n=1: ab = 1440\n", + "\n", + "n=2: (a+d)(b+d) = 1716\n", + "\n", + "n=3: (a+2d)(b+2d) = 1848\n", + "\n", + "Dan P8 = (a+7d)(b+7d)\n", + "\n", + "Sekarang, saya punya:\n", + "\n", + "1) ab = 1440\n", + "\n", + "2) (a+d)(b+d) = 1 \n", + "Extracted:\n", + "1440,\n", + "********************Question:\n", + "Compute the $100^{\\text{th}}$ smallest positive integer $n$ that satisfies the following three congruences:\n", + "\n", + "\\[\n", + "\\begin{aligned}\n", + "\\left\\lfloor \\frac{n}{8} \\right\\rfloor &\\equiv 3 \\pmod{4}, \\\\\n", + "\\left\\lfloor \\frac{n}{32} \\right\\rfloor &\\equiv 2 \\pmod{4}, \\\\\n", + "\\left\\lfloor \\frac{n}{256} \\right\\rfloor &\\equiv 1 \\pmod{4}.\n", + "\\end{aligned}\n", + "\\]\n", + "\n", + "Here, $\\lfloor \\cdot \\rfloor$ denotes the greatest integer function. \n", + "Answer:\n", + "6491 \n", + "Response:\n", + "\n", + "The problem asks for the 100th smallest positive integer n that satisfies three congruences with floor functions. Let me write them down again:\n", + "\n", + "\\[\n", + "\\begin{aligned}\n", + "\\left\\lfloor \\frac{n}{8} \\right\\rfloor &\\equiv 3 \\pmod{4}, \\\\\n", + "\\left\\lfloor \\frac{n}{32} \\right\\rfloor &\\equiv 2 \\pmod{4}, \\\\\n", + "\\left\\lfloor \\frac{n}{256} \\right\\rfloor &\\equiv 1 \\pmod{4}.\n", + "\\end{aligned}\n", + "\\]\n", + "\n", + "First, I need to understand what these congruences mean. The floor function gives the greatest integer less than or equal to n divided by some number, and that integer must satisfy a congruence modulo 4.\n", + "\n", + "Let me denote k = floor(n / 256), since that's the largest divisor.\n", + "\n", + "k is given to be congruent to 1 mod 4, so k = 4m + 1 for some integer m.\n", + "\n", + "n / 256 < k+1, so n < 256(k+1)\n", + "\n", + "And n >= 256k, since floor(n/256) = k.\n", + "\n", + "Now, floor(n/256) = k ≡ 1 mod 4, so k = 4m + 1.\n", + "\n", + "Then n is between 256*(4m+1) and 256*(4m+1) + 255, which is from 1024m + 256 to 1024m + 511? Let me calculate.\n", + "\n", + "256 * (4m + 1) = 1024m + 256\n", + "\n", + "And upper limit is 1024m + 256 + 255 = 1024m + 511\n", + "\n", + "Now, within this range, we have floor(n/256) = k = 4m+1, which satisfies the third congruence: 1 ≡ 1 mod 4.\n", + "\n", + "Now, we also have the second congruence: floor(n/32) ≡ 2 mod 4.\n", + "\n", + "floor(n/32) is approximately n/32, but since n is around 1024m, floor(n/32) should be around 1024m / 32 = 32m.\n", + "\n", + "n min is 1024m + 256, n max is 1024m + 511.\n", + "\n", + "So n/32 min: (1024m + 256)/32 = 32m + 8\n", + "\n", + "1024 / 32 = 32, yes.\n", + "\n", + "256 / 32 = 8, so n/32 min is 32m + 8\n", + "\n", + "n/32 max: (1024m + 511)/32 = 1024m/32 + 511/32 = 32m + 15.96875\n", + "\n", + "Since n is integer, floor(n/32) will be the greatest integer less than or equal to n/32.\n", + "\n", + "n min gives n/32 = 32m + 8, which is integer, so floor is 32m + 8.\n", + "\n", + "n max n/32 < 32m + 16, since 511/32 = 15.96875, so floor(n/32) < 32m + 16.\n", + "\n", + "n/32 max is less than 32m + 16, and since it's less than 32m + 16, floor can be up to 32m + 15.\n", + "\n", + "Now \n", + "Extracted:\n", + "100\n", + "********************Question:\n", + "Compute \\[\\lfloor 1 \\rfloor + \\lfloor 1.6 \\rfloor + \\lfloor 2.2 \\rfloor + \\lfloor 2.8 \\rfloor + \\dots + \\lfloor 99.4 \\rfloor + \\lfloor 100 \\rfloor,\\]where the arguments of the floor functions are in arithmetic progression. \n", + "Answer:\n", + "8317 \n", + "Response:\n", + "\n", + "Saya perlu menghitung jumlah dari serangkaian floor function dengan argumen dalam deret aritmatik. Deretnya dimulai dari 1, kemudian 1.6, 2.2, 2.8, dan seterusnya hingga 100. Saya perlu mengidentifikasi pola dari argumen tersebut.\n", + "\n", + "Pertama, mari lihat bedanya antar suku. Dari 1 ke 1.6, selisihnya adalah 0.6. Lalu dari 1.6 ke 2.2, selisihnya 0.6. Dari 2.2 ke 2.8, selisihnya 0.6. Jadi, ini adalah deret aritmatik dengan suku pertama a = 1 dan beda d = 0.6.\n", + "\n", + "Sekarang, suku terakhir adalah 100. Saya perlu mencari berapa banyak suku dalam deret ini.\n", + "\n", + "Suku ke-n dari deret aritmatik diberikan oleh: U_n = a + (n-1)d.\n", + "\n", + "Saya punya U_k = 100.\n", + "\n", + "Jadi, 1 + (k-1)*0.6 = 100\n", + "\n", + "Kemudian, (k-1)*0.6 = 99\n", + "\n", + "k-1 = 99 / 0.6\n", + "\n", + "99 / 0.6 = 990 / 6, sebaiknya hitung dengan pecahan.\n", + "\n", + "0.6 = 3/5, jadi (k-1) * (3/5) = 99\n", + "\n", + "Maka, k-1 = 99 * (5/3) = (99/3)*5 = 33*5 = 165.\n", + "\n", + "Jadi, k-1 = 165, maka k = 166.\n", + "\n", + "Ada 166 suku.\n", + "\n", + "Sekarang, setiap suku adalah n_i = a + (i-1)d dengan a = 1, d = 0.6.\n", + "\n", + "Tapi floor functionnya, jadi kita perlu floor(n_i) untuk setiap i dari 1 hingga 166.\n", + "\n", + "n_i = 1 + (i-1)*0.6\n", + "\n", + "Mari tulis ulang n_i.\n", + "\n", + "n_i = 1 + (i-1)*0.6 = 1 + (i-1)*(3/5)\n", + "\n", + "Floor(n_i) adalah bilangan bulat terbesar yang kurang dari atau sama dengan n_i.\n", + "\n", + "Sekarang, karena deretnya adalah aritmatik dengan beda 0.6, dan 0.6 < 1, maka floor function akan berubah setiap ketika n_i melewati bilangan bulat.\n", + "\n", + "Saya perlu mengelompokkan suku-suku ini berdasarkan nilai floor-nya.\n", + "\n", + "Pertama, mari cari bilangan bulat mana yang menjadi floor.\n", + "\n", + "Misalnya, untuk n_i < 1, tapi n_i dimulai dari 1, jadi saat i=1, n_1=1.\n", + "\n", + "Floor dari 1 adalah 1.\n", + "\n", + "Saat n_i antara 1 dan 2, floor-nya 1.\n", + "\n", + "Antara 2 dan 3, floor-nya 2, dan seterusnya.\n", + "\n", + "Jadi, perlu mencari ketika suku berada di interval antara bilangan bulat.\n", + "\n", + "Karena beda 0.6, maka dalam satu interval bilangan bulat, misalnya [m, m+1), berapa banyak suku yang ada?\n", + "\n", + "Pertama, kita perlu mencari dari mana deret ini mulai dan berakhir.\n", + "\n", + "n_i = 1 + (i-1)*0.6\n", + "\n", + "Setelah i=1, n_1=1\n", + "\n", + "n_2=1+0.6= \n", + "Extracted:\n", + ".\n", + "********************Question:\n", + "\\[ \\int_{-10}^{10} \\left| 4 - \\left| 3 - \\left| 2 - \\left| 1 - |x| \\right| \\right| \\right| \\right| \\, \\mathrm{d}x \\] \n", + "Answer:\n", + "44 \n", + "Response:\n", + "\n", + "Pertama, saya harus memahami integral yang diberikan. Ini adalah integral tentu dari fungsi mutlak yang bersarang, dari -10 hingga 10. Batasnya simetris, dan karena ada mutlak, mungkin fungsi ini simetris juga.\n", + "\n", + "Fungsinya adalah: | 4 - | 3 - | 2 - | 1 - |x| | |\n", + "\n", + "Saya perlu menyederhanakan ini. Mulai dari dalam, karena mutlak akan mengubah bentuk tergantung pada nilai x.\n", + "\n", + "Mari kita definisikan fungsi langkah demi langkah.\n", + "\n", + "Pertama, | 1 - |x| |. Ini bergantung pada apakah |x| kurang dari 1 atau lebih.\n", + "\n", + "Jika |x| < 1, maka 1 - |x| > 0, jadi | 1 - |x| | = 1 - |x|\n", + "\n", + "Jika |x| > 1, maka 1 - |x| < 0, jadi | 1 - |x| | = |x| - 1\n", + "\n", + "Jadi, | 1 - |x| | = {\n", + " 1 - |x| jika |x| < 1,\n", + " |x| - 1 jika |x| > 1\n", + "}\n", + "\n", + "Tapi karena x dari -10 ke 10, |x| dari 0 sampai 10.\n", + "\n", + "Sekarang, selanjutnya, | 2 - | 1 - |x| | |\n", + "\n", + "Pertama, kita punya | 1 - |x| | yang sudah terdefinisi.\n", + "\n", + "Misalnya, untuk |x| < 1, | 1 - |x| | = 1 - |x|\n", + "\n", + "Kemudian | 2 - (1 - |x|) | = | 2 - 1 + |x| | = |1 + |x||, dan karena 1 + |x| > 0 untuk semua x, jadi |1 + |x|| = 1 + |x|\n", + "\n", + "Tunggu, 2 - | 1 - |x| |, dan pada daerah |x| < 1, | 1 - |x| | = 1 - |x|, jadi 2 - (1 - |x|) = 2 - 1 + |x| = 1 + |x|, dan memang pasti positif, jadi |2 - | 1 - |x| || = 1 + |x| untuk |x| < 1.\n", + "\n", + "Sekarang, jika |x| > 1, | 1 - |x| | = |x| - 1, jadi 2 - (|x| - 1) = 2 - |x| + 1 = 3 - |x|, jadi |2 - | 1 - |x| || = |3 - |x||\n", + "\n", + "Sekarang, |3 - |x|| itu tergantung pada |x|<3 atau >3.\n", + "\n", + "Jadi, mari kita definisikan g(x) = | 2 - | 1 - |x| | |\n", + "\n", + "Untuk |x| < 1:\n", + "g(x) = 1 + |x|\n", + "\n", + "Untuk |x| > 1: \n", + "jika |x| < 3, g(x) = |3 - |x||? Tidak.\n", + "\n", + "Kita bahas kasus.\n", + "\n", + "(g(x)) = | 2 - | 1 - |x| | |\n", + "\n", + "| 1 - |x| | terdefinisi dengan 0.\n", + "\n", + "Poin krusial adalah ketika |x| = 1, dan |x| = 3, dan mungkin dari luar.\n", + "\n", + "Untuk x dari -10 ke 10, |x| dari 0 sampai 10.\n", + "\n", + "Mari kita bag \n", + "Extracted:\n", + ",\n", + "********************Question:\n", + "The number $734,851,474,594,578,436,096$ is equal to $n^6$ for some positive integer $n$. What is the value of $n$? \n", + "Answer:\n", + "3004 \n", + "Response:\n", + "\n", + "Saya diberikan bilangan 734.851.474.594.578.436.096 dan dikatakan itu adalah kuadrat enam dari suatu bilangan bulat positif n. Saya harus menemukan n.\n", + "\n", + "Pertama, saya pikir langkah awalnya adalah mencari akar keen enam dari bilangan ini. Tetapi bilangan seperti itu mungkin besar, jadi saya perlu cara yang bijaksana.\n", + "\n", + "Saya ingat bahwa selain pangkatnya, nilai n mungkin memiliki digit-digit yang enak. Misalnya, n bisa adalah bilangan yang kuadratnya mudah ditemukan atau membagi bilangan tersebut.\n", + "\n", + "Bilangan ini besar, jadi mungkin n adalah pangkat dua atau tiga sendiri, contoh: n = a^b, maka n^6 = a^{6b}, jadi jika b membagi 6, maka mudah.\n", + "\n", + "Tapi tidak ada petunjuk apapun, jadi saya harus mencoba mencari akar manual.\n", + "\n", + "Lihat bilangan lengkapnya: 734851474594578436096.\n", + "\n", + "Mungkin ini adalah kuadrat atau kubik atau sesuatu. Tapi n^6, jadi akar enamnya.\n", + "\n", + "Saya tahu bahwa untuk memeriksa apakah bilangan adalah kuadrat atau kubik, ada cara tertentu.\n", + "\n", + "Tapi n^6 itu sama dengan (n^2)^3 atau (n^3)^2, jadi sebenarnya ini adalah kuadrat dan juga kubik, jadi mungkin mudah untuk mencari akar dua atau tiga.\n", + "\n", + "Akar dua dari n^6 adalah n^3, dan akar tiga adalah n^2, jadi jika saya bisa menemukan akar dua atau akar tiga, lalu akar lagi.\n", + "\n", + "Mungkin lebih mudah untuk mencari akar tiga karena bilangan untuk akar tiga mungkin terlihat lebih kecil.\n", + "\n", + "Tapi bilangan ini sangat besar, jadi akar tiga mungkin masih besar.\n", + "\n", + "Saya pikir ini mungkin adalah bilangan yang kuadratnya atau kubiknya familiar, tapi tidak datang apa-apa.\n", + "\n", + "Mungkin n adalah bilangan yang sendiri punya sifat bagus, seperti bilangan kuasa dua.\n", + "\n", + "Tapi saya tidak punya ide, jadi saya harus mencoba menghitung akar enam dengan calculator. Tapi itu mungkin tidak diizinkan atau tidak praktis dalam konteks.\n", + "\n", + "Mari kita coba hitung digit dari akar.\n", + "\n", + "Bilangan ini punya berapa digit? Mari kita tulis: 734,851,474,594,578,436,2496? Tunggu, ulang: 734,851,474,594,578,436,096.\n", + "\n", + "Mari kita tanam dalam pemahaman saya.\n", + "\n", + "734.851.474.594.578.436.096\n", + "\n", + "Jadi, digitnya: mari hitung digit.\n", + "\n", + "734 (3 digit), 851 (3), 474 (3), 594 (3), 578 (3), 436 (3), 096 (3) jadi 7 kelompok, tapi, tunggu, ketika saya kelompokkan dengan tanda koma, tapi mungkin pengelompokan itu bukan tentang digit, tapi digit string.\n", + "\n", + "Mungkin lebih baik menghitung digit secara langsung.\n", + "\n", + "Bilangan 734851474594578436096\n", + "\n", + "Mari kita \n", + "Extracted:\n", + "734.851.474.594.578.436.096\n", + "********************Question:\n", + "What is the sum of all possible values of $\\cos(2\\theta)$ if $\\cos(2\\theta) = 2\\cos(\\theta)$ for a real number $\\theta$? \n", + "Answer:\n", + "2 \n", + "Response:\n", + "\n", + "Saya perlu mencari jumlah semua nilai mungkin untuk cos(2θ) jika cos(2θ) = 2 cos(θ) untuk suatu angka real θ. Saya akan menyelesaikan persamaan ini untuk menemukan nilai-nilai θ yang memenuhi keduanya, kemudian menemukan cos(2θ) yang mungkin dan menyelesaikan untuk derajat yang diminta.\n", + "\n", + "Pertama, saya ingat identitas ganda untuk cos(2θ): cos(2θ) = 2 cos²(θ) - 1. Jadi, saya bisa menulis persamaan dengan substitusi.\n", + "\n", + "Persamaannya adalah cos(2θ) = 2 cos(θ). Ganti dengan identitas:\n", + "\n", + "2 cos²(θ) - 1 = 2 cos(θ)\n", + "\n", + "Sekarang, biarkan x = cos(θ). Secara aritmetika, ketika θ berubah, x juga berubah, dan x berada di [ -1, 1 ] karena itu cosinus.\n", + "\n", + "Jadi, 2x² - 1 = 2x\n", + "\n", + "Bawa semua suku ke satu sisi untuk membuat persamaan kuadrat.\n", + "\n", + "2x² - 1 - 2x = 0\n", + "\n", + "Jadi, 2x² - 2x - 1 = 0\n", + "\n", + "Sekarang, selesaikan untuk x.\n", + "\n", + "Pakai rumus kuadrat: x = [2 ± √(4 - 4*2*(-1))] / (2*2) = [2 ± √(4 + 8)] / 4 = [2 ± √12] / 4\n", + "\n", + "Buat akarnya lebih baik. √12 = 2√3, jadi x = [2 ± 2√3] / 4 = [1 ± √3] / 2\n", + "\n", + "Jadi, x = [1 + √3]/2 atau x = [1 - √3]/2\n", + "\n", + "Sekarang, saya perlu memeriksa apakah nilai-nilai ini ada di [-1,1].\n", + "\n", + "Pertama, [1 + √3]/2. √3 ≈ 1.732, jadi 1 + 1.732 = 2.732, dibagi 2 adalah 1.366. Tapi cosinus tidak bisa lebih besar dari 1, jadi ini keluar dari rentang. Jadi, hanya x = [1 - √3]/2 yang mungkin.\n", + "\n", + "1 - √3 ≈ 1 - 1.732 = -0.732, dibagi 2 adalah -0.366, yang ada di [-1,1]. Jadi, x = [1 - √3]/2 ≈ -0.366\n", + "\n", + "Tapi x = cos(θ), dan saya punya dua akar, tetapi satu di luar rentang. Jadi, hanya satu nilai x yang mungkin.\n", + "\n", + "Tapi ini harus memberi dua nilai untuk cos(2θ)? Saya pikir saya keliru.\n", + "\n", + "Saya menulis cos(2θ) = 2 cos²(θ) - 1, dan setara dengan 2 cos(θ), jadi saya punya 2 cos²(θ) - 1 = 2 cos(θ)\n", + "\n", + "Tapi ini adalah persamaan untuk cos(2θ), yang sebenarnya adalah eksprisinya dalam variable cos(θ).\n", + "\n", + "Tapi setelah menyelesaikan, saya punya satu nilai x yang mungkin, tapi harusnya ada lebih karena persamaan harus simetris atau apa?\n", + "\n", + "Mungkin saya salah menulis. Persamaannya adalah cos(2θ) = 2 cos(θ)\n", + "\n", + "Saya punya identity cos(2θ) = 2 cos²(θ \n", + "Extracted:\n", + "2\n", + "********************Question:\n", + "An engineer invested $\\$10,\\!000$ in a six-month savings certificate that paid a simple annual interest rate of $12\\%$. After six months, she invested the total value of her investment in another six-month certificate. After six more months, the investment was worth $\\$11,\\!130$. If the annual interest rate of the second certificate is $r\\%,$ then what is $r?$ \n", + "Answer:\n", + "10 \n", + "Response:\n", + "\n", + "Pertama, mari pahami masalahnya. Seorang insinyur menabung $10.000 selama enam bulan dengan suku bunga sederhana tahunan sebesar 12%. Setelah enam bulan, dia memindahkan total investasinya ke sertifikat deposito enam bulan lagi dengan suku bunga sederhana tahunan r%. Setelah enam bulan lagi, nilainya menjadi $11.130. Saya perlu mencari r.\n", + "\n", + "Suku bunga sederhana berarti bunga dibayarkan hanya pada modal awal. Rata-rata bunga tahunan adalah 12%, tapi periode hanya enam bulan.\n", + "\n", + "Pertama, hitung nilai setelah enam bulan pertama.\n", + "\n", + "Suku bunga tahunan 12% untuk enam bulan. Karena sederhana, bunga untuk enam bulan adalah setengah dari tahunan, tapi harus hati-hati.\n", + "\n", + "Suku bunga tahunan 12%, berarti untuk setahun. Untuk enam bulan, yang adalah 0.5 tahun, bunga sederhana adalah (12% / 1) * (0.5) * modal = 0.06 * modal.\n", + "\n", + "Bunga sederhana dihitung sebagai modal awal dikali suku bunga persen dikali waktu.\n", + "\n", + "Disini, suku bunga diberikan sebagai annual, dan waktu dalam tahun.\n", + "\n", + "Jadi, untuk pertama kali investasi: modal awal M0 = $10.000\n", + "\n", + "Suku bunga anual 12% simple interest.\n", + "\n", + "Waktu = 6 bulan = 0.5 tahun.\n", + "\n", + "Bunga untuk enam bulan = M0 * (12/100) * (0.5) = 10000 * 0.12 * 0.5 = 10000 * 0.06 = $600\n", + "\n", + "Bunga sederhana: karena sederhana, bunga hanya pada modal awal, jadi tidak ada kompensasi dari bunga berikutnya.\n", + "\n", + "Jadi, total setelah enam bulan pertama = modal awal + bunga = 10000 + 600 = $10.600\n", + "\n", + "Setelah itu, dia memindahkan ke sertifikat lagi dengan suku bunga r% annual simple interest.\n", + "\n", + "Modal awal untuk yang kedua adalah $10.600\n", + "\n", + "Waktu lagi 6 bulan, atau 0.5 tahun.\n", + "\n", + "Setelah enam bulan kedua, nilainya menjadi $11.130\n", + "\n", + "Jadi, bunga untuk periode kedua adalah $11.130 - $10.600 = $530\n", + "\n", + "Bunga dihitung dengan simple interest, basis modal awal untuk periodenya.\n", + "\n", + "Jadi, bunga = modal awal * suku bunga anual * waktu\n", + "\n", + "Modal awal untuk yang kedua adalah $10.600\n", + "\n", + "Waktu = 0.5 tahun\n", + "\n", + "Suku bunga anual = r% = r/100\n", + "\n", + "Bunga = 10600 * (r/100) * 0.5 = 11.130\n", + "\n", + "Bunga adalah $530, ya.\n", + "\n", + "Jadi, persamaannya:\n", + "\n", + "10600 * (r/100) * 0.5 = 530\n", + "\n", + "Saya dapat menyelesaikannya untuk r.\n", + "\n", + "Pertama, sederhanakan.\n", + "\n", + "10600 * 0.5 * (r/100) = 530\n", + "\n", + "Hitung 10600 * 0.5 = 5300\n", + "\n", + "Jadi, 5300 * (r/100) = 5 \n", + "Extracted:\n", + ",\n", + "********************Question:\n", + "In a far-off land, three fish can be traded for two loaves of bread, and a loaf of bread can be traded for four bags of rice. How many bags of rice is one fish worth? Provide your answer in the form \\(\\frac{k}{m}\\), where \\(k\\) and \\(m\\) are integers, and give the value of \\(k + m\\). \n", + "Answer:\n", + "11 \n", + "Response:\n", + "\n", + "The problem is set in a far-off land with trade rates: three fish for two loaves of bread, and one loaf of bread for four bags of rice. I need to find how many bags of rice one fish is worth, expressed as a fraction k/m, and then find k + m.\n", + "\n", + "First, I need to understand the exchanges. From the first trade, three fish can be traded for two loaves of bread. So, the ratio is fish to bread: 3 fish / 2 bread.\n", + "\n", + "Second, one loaf of bread can be traded for four bags of rice. So, bread to rice: 1 bread / 4 rice.\n", + "\n", + "I need to find the value in terms of rice for one fish.\n", + "\n", + "Since bread is common in both, I should convert everything to bread and then to rice.\n", + "\n", + "Find how many loaves of bread one fish is worth.\n", + "\n", + "From 3 fish = 2 bread, so 1 fish = ? bread.\n", + "\n", + "Set up the equation: 3 fish = 2 bread\n", + "\n", + "So, 1 fish = (2/3) bread.\n", + "\n", + "Okay, so one fish is worth two-thirds of a loaf of bread.\n", + "\n", + "Now, from the other trade, 1 bread = 4 rice.\n", + "\n", + "So, if 1 bread = 4 rice, then (2/3) bread = (2/3) * 4 rice.\n", + "\n", + "2/3 times 4 is 8/3.\n", + "\n", + "So, one fish is worth 8/3 bags of rice.\n", + "\n", + "8/3 is a fraction, so k=8, m=3, and k+m=11.\n", + "\n", + "The problem says to provide answer in form k/m where k and m are integers, which 8/3 is, and I can leave it like that.\n", + "\n", + "But let me double-check.\n", + "\n", + "I could use the rice directly without bread.\n", + "\n", + "Set up a system.\n", + "\n", + "Let F be fish, B be bread, R be rice.\n", + "\n", + "Given 3F = 2B, so B = (3/2)F? No.\n", + "\n", + "From 3F = 2B, so B = (3/2)F? Let's see.\n", + "\n", + "If 3 fish = 2 bread, then one fish = (2/3) bread, so B = (3/2)F only if F =1.\n", + "\n", + "Inverse.\n", + "\n", + "From 3F = 2B, so B = (3/2)F? No.\n", + "\n", + "If 3F = 2B, then for example, if F=3, B=2.\n", + "\n", + "So, B in terms of F: 2B = 3F, so B = (3/2)F? But when F=3, B=2, and (3/2)*3 = 4.5, but B should be 2, which is not 4.5. Mistake.\n", + "\n", + "3F = 2B, so when F=3, B=2.\n", + "\n", + "So, B = (2/3)F * 2? Let's solve.\n", + "\n", + "From 3F = 2B, so B = (3/2)F? 3/2 * 3 = 4.5, but should be 2, which is not.\n", + "\n", + "I think I messed up.\n", + "\n", + "Equation: 3 fish = 2 bread, meaning when you trade 3 fish, you get 2 bread.\n", + "\n", + "So, the number of bread per fish is constant.\n", + "\n", + "So, when you have F fish, B = (2/3) F? Let's say.\n", + "\n", + "If F=3, B=2, so B = (2/3) * F? (2/3)*3=2, yes.\n", + "\n", + "But B is in terms of F.\n", + "\n", + "B = (2/3) F means for example, if F=1, B=2/3.\n", + "\n", + "Now, the second equation: 1 bread = 4 rice.\n", + "\n", + "So, for rice, when you have B bread, R = 4 * B? No.\n", + "\n", + "1 bread can be traded for 4 rice, so when you trade one bread, you get 4 \n", + "Extracted:\n", + ",\n", + "********************Question:\n", + "In triangle $ABC$, $\\sin \\angle A = \\frac{4}{5}$ and $\\angle A < 90^\\circ$. Let $D$ be a point outside triangle $ABC$ such that $\\angle BAD = \\angle DAC$ and $\\angle BDC = 90^\\circ$. Suppose that $AD = 1$ and that $\\frac{BD}{CD} = \\frac{3}{2}$. If $AB + AC$ can be expressed in the form $\\frac{a\\sqrt{b}}{c}$ where $a, b, c$ are pairwise relatively prime integers, find $a + b + c$. \n", + "Answer:\n", + "34 \n", + "Response:\n", + "\n", + "The problem involves a triangle with given angles and conditions. Let me read it carefully.\n", + "\n", + "We have triangle ABC with sin angle A equal to 4/5 and angle A is less than 90 degrees. So, angle A is acute, which makes sense since sin is positive.\n", + "\n", + "First, since sin A = 4/5, and A < 90°, I can find cos A using the Pythagorean identity.\n", + "\n", + "cos²A = 1 - sin²A = 1 - (16/25) = 9/25, so cos A = 3/5 (since angle A is acute).\n", + "\n", + "Okay, so in triangle ABC, angle A is known, with cos A = 3/5, sin A = 4/5.\n", + "\n", + "Now, there's a point D outside the triangle such that angle BAD equals angle DAC, and angle BDC is 90 degrees. Also, AD = 1, and BD/CD = 3/2.\n", + "\n", + "I need to find AB + AC, expressed as a form with square roots, and then sum of coefficients.\n", + "\n", + "First, angle BAD = angle DAC suggests that AD is the angle bisector of angle BAC.\n", + "\n", + "So, AD bisects angle A.\n", + "\n", + "But D is outside the triangle, and angle BDC is 90 degrees, so in triangle BDC, angle at D is 90 degrees, so BD and CD are perpendicular.\n", + "\n", + "AD is the angle bisector, and AD = 1.\n", + "\n", + "Also, BD/CD = 3/2.\n", + "\n", + "Now, I need to find AB + AC.\n", + "\n", + "Since AD is the angle bisector, and we're dealing with lengths, I should use the angle bisector theorem.\n", + "\n", + "But the angle bisector theorem typically applies when D is on BC, but here D is outside the triangle, so probably not.\n", + "\n", + "The angle bisector extends beyond the triangle, and D is on that extension or something.\n", + "\n", + "Let me sketch this.\n", + "\n", + "Let me draw triangle ABC, with angle at A being acute, say with cos A = 3/5, so it's a bit more than 60 degrees, since cos 60° is 0.5, 3/5=0.6, so angle A is about 53 degrees or something, but not necessary.\n", + "\n", + "Now, from A, I draw the angle bisector. Since angle A is acute, the angle bisector will go inside the triangle if B and C are acute, but the problem says D is outside the triangle, so probably the angle bisector is extended beyond C or beyond B.\n", + "\n", + "Now, angle BDC is 90 degrees, so when I go along the angle bisector and hit D such that BD and CD are perpendicular.\n", + "\n", + "Also, BD/CD = 3/2.\n", + "\n", + "Since AD is the angle bisector, and D is on the extension of the angle bisector, I assume.\n", + "\n", + "In standard angle bisector theorem, if D were on BC, then AB/AC = BD/DC, but here D is not on BC, and also BD/CD is given as 3/2, but if D were on BC, BD/CD would be related to the sides.\n", + "\n", + "But here, D is outside, so it's different.\n", + "\n", + "I should consider the properties of the angle bisector and the point D.\n", + "\n", + "Since angle BAD = angle DAC, and AD is common, then in triangles ABD and ACD, but D is not necessarily connected directly.\n", + "\n", + "I could consider the reflection or something.\n", + "\n", + "Since angle BDC is 90 degrees, D lies on the circle with diameter BC, because angle in a semicircle is right angle.\n", + "\n", + "Yes, if angle at D is 90 degrees in triangle BDC, then D lies on the circle with diameter BC.\n", + "\n", + "So, D is on the circle with diameter BC.\n", + "\n", + "Also, AD is the angle bisector, and AD = 1.\n", + "\n", + "Moreover, BD/CD = 3/2.\n", + "\n", + "Now, since D is on the circle with diameter BC, BD and CD are related to the angle subtended.\n", + "\n", + "I could set up \n", + "Extracted:\n", + ".\n", + "********************Question:\n", + "Find the number of $10$-digit palindromes that are not divisible by $11$. \n", + "Answer:\n", + "0 \n", + "Response:\n", + "\n", + "Pertama, soalnya adalah mencari jumlah palindrom 10-digit yang tidak habis dibagi oleh 11. Saya harus berpikir dalam Bahasa Indonesia.\n", + "\n", + "Sebelumnya, saya perlu mengingat apa itu palindrome. Untuk angka 10-digit, bentuknya adalah abcdefghij, tetapi karena palindrome, maka a=j, b=i, c=h, d=g, e=f. Sepertinya untuk 10-digit, indeksnya simetris:\n", + "\n", + "Pertama digit: a, kedua: b, ketiga: c, keempat: d, kelima: e, keenam: f, ketujuh: g, kedelapan: h, kesembilan: i, kelayar: j.\n", + "\n", + "Dalam palindrome, a=j, b=i, c=h, d=g, e=f. Jadi, sebenarnya hanya digit pertama hingga kelima yang bebas, lalu yang tersisa ditentukan secara simetris.\n", + "\n", + "Jadi, untuk 10-digit palindrome, kita memilih digit a, b, c, d, e, dan kemudian:\n", + "\n", + "j = a, i = b, h = c, g = d, f = e.\n", + "\n", + "Jadi, sebenarnya hanya 5 digit yang bebas: a, b, c, d, e. Masing-masing bisa angka 0-9, tetapi karena ini angka 10-digit, a tidak bisa nol. Jadi, a dari 1-9, dan b,c,d,e dari 0-9.\n", + "\n", + "Jadi, total palindrome 10-digit sebenarnya adalah pilihan untuk a,b,c,d,e.\n", + "\n", + "a: 1-9 → 9 pilihan\n", + "\n", + "b: 0-9 → 10 pilihan\n", + "\n", + "c: 0-9 → 10 pilihan\n", + "\n", + "d: 0-9 → 10 pilihan\n", + "\n", + "e: 0-9 → 10 pilihan\n", + "\n", + "Jadi total = 9 * 10 * 10 * 10 * 10 = 90,000\n", + "\n", + "Sekarang, kita ingin yang tidak habis dibagi 11.\n", + "\n", + "Pertama, saya perlu tahu kapan sebuah angka habis dibagi 11.\n", + "\n", + "Kaidah untuk 11: selisih antara jumlah digit ganjil dan genap-nya harus habis dibagi 11, atau hasil mod 11 = 0.\n", + "\n", + "Lebih tepatnya, untuk keprimaan oleh 11, jumlah digit dengan posisi ganjil (dari kiri) dan genap harus memiliki selisih yang habis dibagi 11.\n", + "\n", + "Cara standar: mulai dari digit paling kiri (posisi ganjil), lalu alternatif.\n", + "\n", + "Misal digit d10 d9 d8 d7 d6 d5 d4 d3 d2 d1, dengan d1 paling kanan.\n", + "\n", + "Pertama, tentukan posisi.\n", + "\n", + "Biasanya, kita hitung dari kiri atau kanan? Saya biasa dari kiri, tetapi untuk aturan 11, suatu bilangan habis dibagi 11 jika selisih antara jumlah digit di posisi ganjil dan genap (dengan posisi dimulai dari kiri maupun kanan, konsisten) sama dengan 0 mod 11.\n", + "\n", + "Mari kita atur posisinya.\n", + "\n", + "Anggap digitnya adalah D1 D2 D3 D4 D5 D6 D7 D8 D9 D10, dengan D1 paling kiri, D10 paling kanan.\n", + "\n", + "Aturan keprimaan 11: bilangan N habis dibagi 11 jika (D1 - D2 + D3 - D4 + D5 - D6 + D7 - D8 + D \n", + "Extracted:\n", + ",\n" + ] + }, + { + "name": "stderr", + "output_type": "stream", + "text": [ + "Unrecognized keys in `rope_scaling` for 'rope_type'='yarn': {'attn_factor'}\n" + ] + }, + { + "data": { + "text/plain": [ + "TrainOutput(global_step=100, training_loss=7.181387914911853e-07, metrics={'train_runtime': 4359.8674, 'train_samples_per_second': 0.092, 'train_steps_per_second': 0.023, 'total_flos': 0.0, 'train_loss': 7.181387914911853e-07})" + ] + }, + "execution_count": 22, + "metadata": {}, + "output_type": "execute_result" + } + ], + "source": [ + "# ==================================================\n", + "# 配置并实例化 GRPOTrainer,其中包含若干奖励函数与训练参数。\n", + "# ==================================================\n", + "\n", + "# For optional training + evaluation\n", + "# new_dataset = dataset.train_test_split(test_size = 0.01)\n", + "\n", + "trainer = GRPOTrainer(\n", + " model = model,\n", + " processing_class = tokenizer,\n", + " reward_funcs = [\n", + " match_format_exactly,\n", + " match_format_approximately,\n", + " check_answer,\n", + " check_numbers,\n", + " format_and_language_reward_func,\n", + " ],\n", + " args = training_args,\n", + " train_dataset = dataset,\n", + "\n", + " # For optional training + evaluation\n", + " # train_dataset = new_dataset[\"train\"],\n", + " # eval_dataset = new_dataset[\"test\"],\n", + ")\n", + "trainer.train()" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "id": "tlaUdxC_VHpz" + }, + "source": [ + "\n", + "### 推理\n", + "\n", + "现在测试训练后的模型!首先测试未使用 GRPO 的基线模型:" + ] + }, + { + "cell_type": "code", + "execution_count": 23, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/", + "height": 190, + "referenced_widgets": [ + "4a69932266bc4a46bf72764abbe0bacb", + "e9c1cebddffd4753a91c895a9ef8ebb4", + "e25cf2e1e05a425bae6330c71774efcf", + "f4fc91122392442da57ec2862cb3d90d", + "2fa427504316456abd55795cbf15f96a", + "8b5a4d10321446fb869172f7165c7cd7", + "fa13a14340074376a0fa8d14e25ed4ba", + "f08f885e86bb43cd960878dbbc35b99f", + "fcb1520bc20e40e8b6cd276fd6892f75", + "b1c1ec8f20a4497da3988ace5424f6fd", + "76534d6f990243f1a635665d2e47a33b" + ] + }, + "id": "qtcz_lpbVC92", + "outputId": "d9fb214a-a516-47d7-ce73-ef9290f74fc9" + }, + "outputs": [ + { + "data": { + "application/vnd.jupyter.widget-view+json": { + "model_id": "adfe16b9d2494f988fe7b9c784e48a7c", + "version_major": 2, + "version_minor": 0 + }, + "text/plain": [ + "Processed prompts: 0%| | 0/1 [00:00 Calculus (Single Variable), Free Online Books\\nWhat is the sqrt of 101?\\nadmin June 23, 2022 Leave a comment\\nsqrt(101) = ? Here we can see the solution step by step.\\nFirst, find out what 10^2 and 11^2 gives.\\n10^2 = 100\\n11^2 = 121\\nSo, 10 and 11 are the two numbers. And 100 and 121 are the squares.\\nBut 100 is closer to 101, so sqrt(101) is approximately 10.something.\\n\\nSo, let me do the calculation for the square root of 101.\\n\\nLet me take a=10, b=101. Then, sqrt(101) = sqrt(b) = a + (b - a^2)/(2a) for initial approximation.\\n\\nBut a^2 is 100, and b is 101, so (b - a^2)/(2a) = (101-100)/20 = 1/20 = 0.05\\n\\nSo, sqrt(101) ≈ 10 + 0.05 = 10.05\\n\\nBut that can’t be right because 10.05^2 = 10.05 * 10.05.\\n\\nDo 10.05^2.\\n\\n10^2 = 100\\nBut 10.05 * 10.05.\\n\\n(10 + 0.05)^2 = 100 + 2*10*0.05 + (0.05)^2 = 100 + 1 + 0.0025 = 101.0025\\n\\nOh, so that gives 101.0025, which is very close to 101. The error is only 0.0025, so actually 10.05 is very close.\\n\\nBut is 10.05 exactly the square root? Not exactly, because 10.05^2=101.0025, which is slightly more than 101.\\n\\nSo sqrt(101) is slightly less than 10.05.\\n\\nBut how much less? Let me try the next method.\\n\\nI can use the fraction method.\\n\\nAnother way to calculate sqrt(101).\\n\\nPerhaps I can use the formula again but with a better approximation.\\n\\nSince 10.05^2 = 101.0025, then sqrt(101) = sqrt(101.0025 - 0.0025) ≈ 10.05 - (0.0025)/(2*10.05)\\n\\nBecause derivative of x^2 is 2x, so for small change.\\n\\nSo, x^2 = 101.0025, dx^2 = -0.0025, then dx/dx^2 = 1/(2x)\\n\\nSo, x = 10.05, dx = dx^2 * (1/(2x)) = -0.0025 / (2*10.05)\\n\\nFirst, 2*10.05=20.10\\n\\nAnd -0.0025 / 20.10 = ? 0.0025 / 20.10 = 0.000124378, so approximately -0.000124\\n\\nSo, sqrt(101) ≈ 10.05 - 0.000124 = 10.049876\\n\\nNow, check 10.049876^2.\\n\\nSince we're very close, or perhaps I can use a calculator, but since this is a math exercise, let's see.\\n\\nThere's another way using continued fractions or something, but that might be complicated.\\n\\nPerhaps 101 is a prime number, so sqrt(\"" + ] + }, + "execution_count": 23, + "metadata": {}, + "output_type": "execute_result" + } + ], + "source": [ + "# ==================================================\n", + "# 使用 fast_generate 测试模型在单条输入上的推理输出。\n", + "# ==================================================\n", + "\n", + "text = \"What is the sqrt of 101?\"\n", + "\n", + "from vllm import SamplingParams\n", + "sampling_params = SamplingParams(\n", + " temperature = 1.0,\n", + " top_k = 50,\n", + " max_tokens = 1024,\n", + ")\n", + "output = model.fast_generate(\n", + " [text],\n", + " sampling_params = sampling_params,\n", + " lora_request = None,\n", + ")[0].outputs[0].text\n", + "\n", + "output" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "id": "Colxz9TAVMsi" + }, + "source": [ + "现在使用我们刚用 GRPO 训练得到的 LoRA,首先保存 LoRA 权重:" + ] + }, + { + "cell_type": "code", + "execution_count": 24, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/" + }, + "id": "AL-BcuB1VLIv", + "outputId": "c23a911b-57b8-4af6-cce4-8cfc3793a856" + }, + "outputs": [ + { + "name": "stderr", + "output_type": "stream", + "text": [ + "Unrecognized keys in `rope_scaling` for 'rope_type'='yarn': {'attn_factor'}\n" + ] + } + ], + "source": [ + "# ==================================================\n", + "# 保存 LoRA 适配器权重到 'grpo_lora' 目录。\n", + "# ==================================================\n", + "\n", + "model.save_lora(\"grpo_lora\")" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "id": "a4LMOBl8boGX" + }, + "source": [ + "验证 LoRA 是否确实被训练:" + ] + }, + { + "cell_type": "code", + "execution_count": 25, + "metadata": { + "id": "4SfdI-ERbpiw" + }, + "outputs": [], + "source": [ + "# ==================================================\n", + "# 使用 safetensors 检查 LoRA 权重文件中的张量是否成功训练(非全零)。\n", + "# ==================================================\n", + "\n", + "from safetensors import safe_open\n", + "\n", + "tensors = {}\n", + "with safe_open(\"grpo_lora/adapter_model.safetensors\", framework = \"pt\") as f:\n", + " # Verify both A and B are non zero\n", + " for key in f.keys():\n", + " tensor = f.get_tensor(key)\n", + " n_zeros = (tensor == 0).sum() / tensor.numel()\n", + " assert(n_zeros.item() != tensor.numel())" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "id": "CwpbwnDBVRLg" + }, + "source": [ + "加载 LoRA 并进行测试,此处不使用自定义 system prompt,以尽量保持模型原有的推理能力:" + ] + }, + { + "cell_type": "code", + "execution_count": 26, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/", + "height": 190, + "referenced_widgets": [ + "4e1e6351f86a4cf99ea07d6a1dd95b3e", + "47a7688ba9044ba29ff56f32244b416d", + "935cb5ca7b2f45f0884a59e3b9f3a7f4", + "d324dd4ef84047aaaf46c5ccb9607638", + "7f52de137c0748629d891d0e6223b81d", + "1fe15cce61e944d58440251cc97946de", + "7342685a5c6b4d2f8b69958792705d30", + "a62fed49c7a14e14811ce4dfa4e02771", + "4b7ef318e5a444a4808910bdd24699c8", + "eb790732f82f4973a2669eecc057f711", + "18f1fe2f63134275a131896d4fe1eb6d" + ] + }, + "id": "X6lXk47v1O4b", + "outputId": "c0ddc7c8-fefa-4883-a128-304da0fec023" + }, + "outputs": [ + { + "data": { + "application/vnd.jupyter.widget-view+json": { + "model_id": "7780076abf9c4010a5cd7e6f57004c3c", + "version_major": 2, + "version_minor": 0 + }, + "text/plain": [ + "Processed prompts: 0%| | 0/1 [00:00\\nI have this equation to solve: (x + 2)^2 = 0. I need to find the value of x that makes this true.\\n\\nThis equation says that (x + 2) squared equals zero. Since it\\'s a square equal to zero, I know that only zero raised to any power is zero. So, for (something)^2 to be zero, that \"something\" must be zero.\\n\\nIn this case, the something is (x + 2), so x + 2 must be zero.\\n\\nTherefore, x + 2 = 0.\\n\\nNow, solving for x, subtract 2 from both sides: x = -2.\\n\\nI can think about this as the equation (x + 2)^2 = 0, which means x + 2 = 0 because if a square is zero, the thing being squared must be zero.\\n\\nIf I expand this, (x + 2)^2 = x^2 + 4x + 4, and set that equal to zero: x^2 + 4x + 4 = 0.\\n\\nNow, factor this quadratic. I know it factors to (x + 2)(x + 2) = 0, or (x + 2)^2 = 0.\\n\\nSo again, x + 2 = 0, so x = -2.\\n\\nUsing the quadratic formula: x = [-b ± sqrt(b^2 - 4ac)] / 2a, where a = 1, b = 4, c = 4.\\n\\nDiscriminants: b^2 - 4ac = 16 - 16 = 0, so sqrt(0) = 0, so x = [-4 ± 0]/2 = -4/2 = -2.\\n\\nAgain, same answer.\\n\\nSince it\\'s a square, there\\'s only one solution, not two like in other cases with negative discriminants.\\n\\nSo, x = -2.\\n\\nThe equation is (x + 2)^2 = 0, which is satisfied only when x + 2 = 0.\\n\\nIn terms of graphing, it would be a parabola touching the x-axis at (-2, 0), so only one root.\\n\\nI think I\\'m confident with this.\\n\\nThe problem is to solve it, so I should just state the answer.\\n\\nBut let me write it properly.\\n\\nThe solution is x = -2.\\n\\nAnd it\\'s a single solution.\\n\\nSometimes people might think to take square root both sides, but taking square root of both sides, I get |x + 2| = 0, which means x + 2 must be zero.\\n\\nYes.\\n\\nSo, I think that\\'s it.\\n\\nThe equation \\\\((x + 2)^2 = 0\\\\) requires that \\\\(x + 2 = 0\\\\), since a square is equal to zero only when the input is zero.\\n\\nSolving for \\\\(x\\\\): \\n\\\\(x + 2 = 0\\\\) \\n\\\\(x = -2\\\\)\\n\\nThis is the only solution, as raising any non-zero value to the power of 2 results in a positive number, not zero.\\n\\n\\\\boxed{-2}'" + ] + }, + "execution_count": 26, + "metadata": {}, + "output_type": "execute_result" + } + ], + "source": [ + "# ==================================================\n", + "# 示例:不带 system prompt,使用 LoRA 进行推理。\n", + "# ==================================================\n", + "\n", + "messages = [\n", + " {\"role\": \"user\", \"content\": \"Solve (x + 2)^2 = 0\"},\n", + "]\n", + "\n", + "text = tokenizer.apply_chat_template(\n", + " messages,\n", + " add_generation_prompt = True, # Must add for generation\n", + " tokenize = False,\n", + ")\n", + "from vllm import SamplingParams\n", + "sampling_params = SamplingParams(\n", + " temperature = 1.0,\n", + " top_k = 50,\n", + " max_tokens = 2048,\n", + ")\n", + "output = model.fast_generate(\n", + " text,\n", + " sampling_params = sampling_params,\n", + " lora_request = model.load_lora(\"grpo_lora\"),\n", + ")[0].outputs[0].text\n", + "\n", + "output" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "id": "g399AC2B1O4b" + }, + "source": [ + "接下来使用带有新语言 system prompt 的方式进行测试:" + ] + }, + { + "cell_type": "code", + "execution_count": 27, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/", + "height": 172, + "referenced_widgets": [ + "622c7d8a89964ee3871f60d1d32dd9d7", + "74655a758460452283c5487e67a09f98", + "518811664352462bbd70d54d7ff75836", + "a45642e387ce4d0a8cf69e1f36d0b279", + "09aa5ac028274fecb5b8a777ea7709c2", + "c937e93a5bc44a639fd4f7f543f86513", + "46e960f31b954d779413192779825fdd", + "872b861723d5466b99002d880af2987e", + "39ad0bd8a1d74242bd72404859d7ca27", + "54fb6e96cfeb42b7a64afa1dcecb080a", + "34ffed32114e410a964ba6bc9f4444fa" + ] + }, + "id": "zf_OY5WMVOxF", + "outputId": "6522d0ea-e766-4964-f0b1-bf1fe0631867" + }, + "outputs": [ + { + "data": { + "application/vnd.jupyter.widget-view+json": { + "model_id": "c842b186d990487aa2fdef786ee139b9", + "version_major": 2, + "version_minor": 0 + }, + "text/plain": [ + "Processed prompts: 0%| | 0/1 [00:00\\nOkay, user wants me to solve (x + 2)^2 = 0. But they also specified that I must process this in Indonesian. Interesting constraint.\\n\\nFirst, let\\'s analyze the equation. This is a simple quadratic where x+2 is squared and equals zero. I remember that squaring any number gives zero only if that number is zero itself. So (x+2) must be zero.\\n\\nBut wait, the user said this must be done in Indonesian. They want my thought process in Bahasa Indonesia. Hmm, I need to switch all internal reasoning to Indonesian temporarily to satisfy the constraint. But the math logic will be the same.\\n\\nBreaking it down: the equation is a perfect square set to zero. The solution is straightforward. I should explain the reasoning step by step since the user asked for \"working out\". \\n\\nFirst, we can take square root both sides. Remember, square root gives both positive and negative options? Wait no, since square root is defined to be non-negative, but in equations we consider both positive and negative solutions.\\n\\nWait, actually: sqrt(a^2) = |a|, but in solving equations we consider both positive and negative for the base case. Let me recast it properly.\\n\\nActually simpler: if a^2=0 then a must be 0. So (x+2)^2=0 implies x+2=0. Then x=-2.\\n\\nBut wait, is this correct? Let me verify. If x=-2, then (-2+2)^2 = 0^2=0. Perfect.\\n\\nIn Indonesian, the term \"persamaan kuadrat\" means quadratic equation, and \"akar kuadrat\" means square root. The solution would be x = -2, and we can say \"dengan demikian solusi yang memenuhi persamaan tersebut adalah x = -2\".\\n\\nThe user only provided one equation though. This is as simple as it gets. But they might want confirmation about why this is the only solution.\\n\\nI recall that for any real number, a number squared is only zero if the number is zero. So indeed only one solution exists here.\\n\\nAlso, in complex numbers it would be the same because |z|^2 = z * conjugate(z) = 0 implies z=0. But user didn\\'t specify context, so probably real numbers are assumed.\\n\\nThe answer is simply x = -2, and in Indonesian \"x di sini adalah -2 atau sebenarnya karena hanya satu solusi\".\\n\\nI think that\\'s it. Simple quadratic equation with unique solution.\\n\\nPertama, mari perhatikan persamaannya: (x + 2)^2 = 0.\\n\\nKita tahu bahwa bilangan kuadrat (dikali dengan dirinya sendiri) hanya akan bernilai nol jika bilangannya tersebut adalah nol. Jadi, dari persamaan (x + 2)^2 = 0, maka:\\n\\n(x + 2)^2 = 0 \\n⇒ x + 2 = 0 (karena hasil kuadrat dari suatu bilangan adalah nol hanya ketika bilangan tersebut adalah nol)\\n\\nSehingga, x = -2.\\n\\nDengan demikian, solusi yang memenuhi persamaan tersebut adalah x = -2.'" + ] + }, + "execution_count": 27, + "metadata": {}, + "output_type": "execute_result" + } + ], + "source": [ + "# ==================================================\n", + "# 示例:带 system prompt,但不使用 LoRA 进行推理。\n", + "# ==================================================\n", + "\n", + "messages = [\n", + " {\"role\": \"system\", \"content\": system_prompt},\n", + " {\"role\": \"user\", \"content\": \"Solve (x + 2)^2 = 0\"},\n", + "]\n", + "\n", + "text = tokenizer.apply_chat_template(\n", + " messages,\n", + " add_generation_prompt = True, # Must add for generation\n", + " tokenize = False,\n", + ")\n", + "from vllm import SamplingParams\n", + "sampling_params = SamplingParams(\n", + " temperature = 1.0,\n", + " top_k = 50,\n", + " max_tokens = 2048,\n", + ")\n", + "output = model.fast_generate(\n", + " text,\n", + " sampling_params = sampling_params,\n", + " lora_request = model.load_lora(\"grpo_lora\"),\n", + ")[0].outputs[0].text\n", + "\n", + "output" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "id": "Ad5qCZMsW_Ed" + }, + "source": [ + "对比带 system prompt 但不加 LoRA 的结果:" + ] + }, + { + "cell_type": "code", + "execution_count": 28, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/", + "height": 172, + "referenced_widgets": [ + "f35b16000bb448139b0954c4eeff99f8", + "ac02d83c54314dd889853d5dfbfcae6f", + "ccb09cc19f804fbca9e3976116227460", + "efa40eb719fd472f804af545734a4fc0", + "17813750f23d44fb9c9305506030709f", + "0693028d222b43e9852beec995ecb5dc", + "7fecf7366e9d41cca37f32d736eed6e6", + "783d83041ef2494287de1352d66e7e85", + "59b5cb2fc77d46d395b22ef63954a7f4", + "3b161341e8d64f5392ebcebf8629df1e", + "e98eff347291445f851d14470b0552cb" + ] + }, + "id": "ee10WWhDW_Ee", + "outputId": "3f7e2b03-c562-46b7-f9d3-27a36af3b473" + }, + "outputs": [ + { + "data": { + "application/vnd.jupyter.widget-view+json": { + "model_id": "c18cad45925648a8a3986934051f5a27", + "version_major": 2, + "version_minor": 0 + }, + "text/plain": [ + "Processed prompts: 0%| | 0/1 [00:00\\nSaya perlu menyelesaikan persamaan (x + 2)^2 = 0. Persamaan ini terlihat sederhana. Karena ini adalah persamaan kuadrat, tetapi disajikan dalam bentuk kuadrat sempurna.\\n\\nSaya ingat bahwa (a + b)^2 = a^2 + 2ab + b^2, tapi di sini, (x + 2)^2 = 0. Jadi, mari saya selesaikan dengan mencari nilai x yang memenuhi persamaan tersebut.\\n\\nJika (x + 2)^2 = 0, maka x + 2 harus sama dengan 0, karena hanya bilangannya sendiri yang mengalikan dirinya sendiri hasilnya nol, yaitu angka nol itu sendiri.\\n\\nJadi, x + 2 = 0. Kemudian, x = -2.\\n\\nSaya bisa memeriksa jawaban ini. Substitusi x = -2 ke dalam persamaan.\\n\\n(x + 2)^2 = (-2 + 2)^2 = (0)^2 = 0, ya, memang benar.\\n\\nApakah ada solusi lain? Dalam persamaan kuadrat biasanya ada dua solusi, tapi di sini hanya satu solusi karena ini adalah bentuk (something)^2 = 0, yang hanya memiliki satu akar.\\n\\nSaya tahu itu sifat akar kuadrat, bahwa jika sesuatu kuadrat sama dengan nol, maka sesuatu itu harus nol.\\n\\nJadi, tunggal solusinya.\\n\\nPertanyaannya hanya meminta untuk menyelesaikan, tidak meminta untuk menemukan akar kuadrat atau sesuatu, hanya persamaan ini.\\n\\nJadi, jawabannya x = -2.\\n\\nSekarang, karena user meminta untuk berpikir dalam Bahasa Indonesia, saya harus menjawab dalam Bahasa Indonesia.\\n\\nTapi pertama, saya telah menyelesaikan.\\n\\nSaya akan menyajikan jawabannya dan kemudian penjelasannya dalam Bahasa Indonesia.\\n\\nPertanyaannya dalam bahasa Inggris, tapi user mengatakan \"Think about the problem and provide your working out. You must think in Bahasa Indonesia.\" Jadi, saya harus memikirkan prosesnya dalam Bahasa Indonesia, tetapi karena ini jawaban akhir, mungkin harus dalam bahasa Inggris atau sesuai konteks.\\n\\nLihat pesan user: \"You must think in Bahasa Indonesia.\" Tapi pesan ini dalam bahasa Inggris. Mungkin maksudnya jawabannya harus dalam Bahasa Indonesia.\\n\\nSaya akan coba jawab dalam Bahasa Indonesia.\\n\\nPertama, mari saya tuliskan prosesnya.\\n\\nPerhatikan bahwa persamaan (x + 2)^2 = 0.\\n\\nKarena kuadrat suatu bilangan hanya bernilai nol jika bilangan tersebut adalah nol, maka:\\n\\nx + 2 = 0\\n\\nSehingga, x = -2.\\n\\nMungkin ada beberapa cara, tapi ini adalah cara yang paling mudah.\\n\\nSaya bisa memanfaatkan sifat akar kuadrat. Perkalian dua bilangan positif atau dua bilangan negatif menghasilkan bilangan positif, dan nol adalah satu-satunya bilangan yang ketika dikalikan dengan dirinya sendiri hasilnya nol.\\n\\nTapi itu lebih dari yang diperlukan.\\n\\nJadi, solusinya adalah x = -2.\\n\\nSekarang, tentang format jawaban. Karena user meminta untuk menyelesaikan, saya akan memberikan jawabannya.\\n\\nNamun, karena itu meminta untuk berpikir, mungkin saya perlu menulis prosesnya secara lengkap.\\n\\nDalam konteks ini, saya akan menulis:\\n\\nPertama-tama, persamaan kuadrat (x + 2)^2 = 0.\\n\\nAmbil akar kuadrat dari kedua ruas, tapi dalam hal penyelesaian, lebih mudah jika memfaktorkan atau menyelesaikan langsung.\\n\\nKarena (x + 2)^2 = 0, maka x + 2 = 0 (karena akar kuadrat dari nol adalah nol).\\n\\nLalu, x = -2.\\n\\nUntuk memeriksa, ganti x dengan -2: ( -2 + 2 )^2 = 0^2 = 0, sehingga benar memenuhi persamaan.\\n\\n'" + ] + }, + "execution_count": 28, + "metadata": {}, + "output_type": "execute_result" + } + ], + "source": [ + "# ==================================================\n", + "# 示例:带 system prompt 并加载 LoRA 后进行推理。\n", + "# ==================================================\n", + "\n", + "messages = [\n", + " {\"role\": \"system\", \"content\": system_prompt},\n", + " {\"role\": \"user\", \"content\": \"Solve (x + 2)^2 = 0\"},\n", + "]\n", + "\n", + "text = tokenizer.apply_chat_template(\n", + " messages,\n", + " add_generation_prompt = True, # Must add for generation\n", + " tokenize = False,\n", + ")\n", + "from vllm import SamplingParams\n", + "sampling_params = SamplingParams(\n", + " temperature = 1.0,\n", + " top_k = 50,\n", + " max_tokens = 2048,\n", + ")\n", + "output = model.fast_generate(\n", + " text,\n", + " sampling_params = sampling_params,\n", + " lora_request = None,\n", + ")[0].outputs[0].text\n", + "\n", + "output" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "id": "EYqpfCF0W_Ee" + }, + "source": [ + "从数据集中抽取 20 个样本,比较使用 LoRA 与不使用 LoRA 的正确语言输出数量:" + ] + }, + { + "cell_type": "code", + "execution_count": 29, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/" + }, + "id": "mJmztPHdW_Ef", + "outputId": "2ea0c34d-abea-45c2-b04b-616323f9f1aa" + }, + "outputs": [ + { + "data": { + "text/plain": [ + "Dataset({\n", + " features: ['prompt', 'solution', 'data_source', 'source_prompt', 'ability', 'reward_model', 'extra_info', 'answer'],\n", + " num_rows: 20\n", + "})" + ] + }, + "execution_count": 29, + "metadata": {}, + "output_type": "execute_result" + } + ], + "source": [ + "# ==================================================\n", + "# 随机抽取 20 条样本组成 sample_dataset 供后续评估。\n", + "# ==================================================\n", + "\n", + "sample_dataset = dataset.shuffle(seed = 3407).select(range(20))\n", + "sample_dataset" + ] + }, + { + "cell_type": "code", + "execution_count": 30, + "metadata": { + "colab": { + "base_uri": "https://localhost:8080/", + "height": 1000, + "referenced_widgets": [ + "47fe26074e4748109e0e3e3b5e437da0", + "f6b16c88dec2470783aad9327b744a30", + "346c6c864b5a47d3aafed9779efe8e7e", + "17a278ccd27b455d8d4c1ef9217720a1", + "bdcc0fe600254330b0d3b26ffdf54e88", + "f9d137906cd34198a886ec03d29e1d25", + "ace6a3ba68d34ae0b402461d0e7fd0ce", + "ddae813db1f445b39b87d5e37d927258", + "f2cab422ad1745cc9e2a57f2b7cbe6a5", + "54ad84288d284fb68fe725a99123d078", + "09e0d1738e5047969b1b53a8051000b6", + "102d1dc21059495c8987fd5473187bda", + "a79a61e313d34c138725d5a36e22b498", + "43dfed9b87664340847c48d5fbc54a7b", + "4ae3ec784a2d42c5805ff73dd3c4e599", + "f8f86f755c0940cf99bec24363177e30", + "283fea36af804c01b7281548570c1fcf", + "c530f911290f4b6abf7c80dcd221d14b", + "0e3a5b3df1c3483f9b49968f2737df07", + "9078876536734ef8a2554b4741b3583c", + "e184cfe6204446edac5eebb249ee5a6c", + "362fb4499d2349ff9700d484f0d75859", + "d863ad0a6a16440daf5c5dab3c2ae8df", + "a77f90d13e6c4777bbcc967ef2d867f5", + "6d352bbed70a4487a0875fc5b53cb417", + "6e7b2a041abf4f83afed2670e1b7cd64", + "0a13e3da3fba4407960e565c4337f407", + "c5a0d96ea1de433ab152dd828bd5206d", + "758d89d3d4ea4084a18bc13216789357", + "8f78548a239f4ad797233826ca2d18a0", + "437b9000fdb64f5685a63f3cf062c6b4", + "3fe86e05b86545ccaceb6b44f7a646f4", + "76658f05feaf46d4afd4d21445a98b46", + "e35c0be4a7bb417f9c79f4195bfe8b80", + "aeb8bbb7f8de402e9db1bc662115ae7c", + "04762bc620b44de1828f231355ecdd90", + "c373be24c8644d21bfbfde16d495f2ec", + "a184d9c3f4be47fc99d67a24e43c6fca", + "a3e84ecd74124797a271ea1bce32e162", + "e29fcc716fcf40929c24cdc89a7a9187", + "1db747fb25f7480ea51477b12c06b2f3", + "5f558701f9b843b9b88deb238006a1d8", + "6d69af9e759a4d7e9ea8dc0a444a2c15", + "a1cb83589109491792b38e2e977154cc", + "c8da96ee22bb467ab6c3b2a4c56c600b", + "ee41fc97fb0f47cbb5abe3e995a18409", + "1a870f1827594b5a96fd4d155ca748f2", + "2082e29d6b5e461a8b53c75a9493acd1", + "200c202ce15545c28791128def355f7f", + "7ffc6b43a2204306933c9dc3e79f73a8", + "53f260ad83924d5f936639da54e3b711", + "ae5c9c693c0b4d3a9390abc0b69d1b09", + "669d68cc8f5d4922aa50acf9436b91a7", + "45bad1acdb5347c9b7bcc16aa33022f5", + "19646a2143d246218fdf3527bbaa4944", + "835635745d8444c094328b88227f83f0", + "3b0d818371de4b788ae5d301f3b77d40", + "72c3038d15584fd186acd123a8f1a70a", + "94753efe1c3d4e9580d59a7a64159455", + "7392747572734c958e3509093f5232ee", + "7bc7fb6cdf0042f597157b85f759069f", + "cc61f72a69e44703b430a7b4bfac5d11", + "2659dab28e8f4ce5931ff109bba50752", + "d87e263201424b0c91afe1ac54572e46", + "f565b58dd1e745208f6b56dd14e7cba3", + "0ea5309034cb45ef854fa9d43da0c9a8", + "3f8078fc9a3b4ba58c01596d280de632", + "c867662d81734b07bc41ae3b6d989fbd", + "5ff2f119429e43af958b34862cf88ce3", + "4b946e624f534c87b52844fbda1c22c5", + "0cc3177fc6234ac3b19bfbada11976fa", + "ae6afa77cfd94387b234842e7b66e411", + "8dd9ea1ecf464c048d6b0620187b2adb", + "4a498b8bc6ff41619ab51f7ddf444f21", + "738fbce122464c9c9ae2543b6a9c6221", + "98d32f0666504cab8e3eb07647170931", + "492fc78e7c29433b93e3380a62a36fd4", + "622a9c4668ed4ce496eca13544ebd68e", + "43cf985162de4a23b8c3edd514a2dc0c", + "4ed61277a1cf4b83934a8f6270291e33", + "ae5c4e6818324bfab094faae4a36863b", + "d57a3a971c7140529aa0219fe755b771", + "f235d4c09a924bdba005c1455eba2a5d", + "c553b7b45bfa44c4b733c9951e33de49", + "c4077902dd614314af0319e82ec8f22e", + "0c9419519ab24681824826f94ff83460", + "c3a7975fd0ea4bfea2522beac4ebe11c", + "ff0a956721284a99ba368a9de27e93e2", + "6815db6af3634835b02c176b54e49286", + "4ef815b8a6a2446aa64cd6f6f23f1cec", + "a8ac6bda67b74bf8bb453825dcc81504", + "520e49e0ae914a568e7f472853b83263", + "ee9905818b104b41af14c30cb62947f8", + "b12de3a1524749d59edeea1d46bfdf3d", + "015819f39cc845d58c5cdf14b52796f8", + "68edf6e2eb064339befe44d77ebe5f19", + "fc1234420dca4ff4a0a7fd3bf39d9cc5", + "a95bc1bba4614346906d22679e0a5700", + "21da79535a5e4715b8d0bbe6a84c4156", + "b5b55ae398044f9797ad9182f0ccc823", + "23d38bbeef094de7b54440ff122b2627", + "c714f65b6a7740e4bd597e941f0c294b", + "7d6f0dd934ce43a6b319e37f37789fd2", + "3fc498ccb014473ba816a42b39d0119b", + "1ac39370c6084132989dd868b7454345", + "2156d69b6028433d8e00edd98b5081f7", + "a57a955d56e64cefb39c6c9c0674c9c4", + "39244d0a8b404a07b3aac7dbbeab4ae2", + "eb59debe975546eaa3b6790cce0cf22f", + "41573bf3e7a74f67a5e923b3bf320cf0", + "7d456bcab176487bb92d7945aa208985", + "8164f92e06e5496081ebc95c3bc5bcfc", + "9c5a549bfaa14591b64a7ee0c508e787", + "d5306ecc83aa46f4b9b32d83f6ea39a5", + "573e549f4c9a4bccad7785d2dca0d085", + "6280bcf7c44d40d3ae7fd5a55b4435f1", + "19746b5c9027445fa337e45c65f88b18", + "c46943f6bdc64cfd9aeb0a54c2814aba", + "5c4c43f02b1d4f2caad6bc2d0c972ca4", + "e6f13effff81490f971915386d700b28", + "4a4deadbd74b4607ad2ebbe67cd943a4", + "2196b8cd440e478784de47e5dc099275", + "e75ec870cdf04e85ac6bead094472d83", + "40c112a0b0434dad8f1986d360a64f67", + "c4e47c48c38f40eb9bd9f15da84dfb25", + "b6cc8cbb2e2a4eb389b38a38d59ca746", + "36deae094c444c52a2d41ccebcba30d9", + "72e45e105b1a4a66bbaefcb2b99244a8", + "e57502b7af9341cea46d2f319df6dc84", + "ca3c7377cfca4defb46ce81e4252056b", + "7bf5b846f29641a1ae06a001a6566af0", + "ba3d68f9730f4531931440e82fc20259", + "4cebd3a225554c0a8e0c4a5323e5b5f3", + "c68eabf0e9084148af917ad69d0f291a", + "11dca6b14f654a3f881d7942fcb76b11", + "730bd5d31b5c469e9d01103a61044b5d", + "76bef67505e447d79bad5d759d5ae5c6", + "25cf2302cb954727b2a7fa76c9a3e559", + "d3ea78e924944f1eb399d6a0fc46bba3", + "d9b0563fa75046859b311a6a6b924754", + "52ffffe2447f4d46851cf2091d2489a5", + "c1dcdfff861649ce83651e1eff1a3c49", + "971653b364ea4526869e4f451ac4f1e3", + "b1ed69dc1a2247f7bb088578a98c9c79", + "036640f29e2d4f92aebd5e155b19644f", + "d37d413784864f008b6d90701d2f158b", + "2ffb26fcb39b4faf95680890dff651ad", + "91e6b34c909a47dda1486851524aa258", + "ffbf7936ce484cfcaf0648fcfc920286", + "d5fd8382c3104e9badfd19deb219c2fa", + "a2ce361e5e324aa084ad7d171e7c6cb6", + "d0e3c9f8274a43649ac0f8edaea63788", + "19da85bb1cc64ca286eba0b8d53914ed", + "395e110f29b34500804a0bdffddef7dd", + "b2dd5fd61aaf4235b9011856df2b0f46", + "c162fd2ae9274ef691982e22f4382054", + "2824756c51aa46a3a80201be5c51700b", + "b479c1a1c21f42279350f1ef46c96b84", + "24ce824a8bdc44979717dc8fc6bdc2b1", + "4964a52db09e4e7b8ab80f10e6063fa1", + "aad966f64273400ea4c947a51da9a81f", + "db3576555492431bb423feba4f76b8d0", + "1f86b4dfe21e47ada984ac313efa0817", + "136529877dbe4bb5a814c774e629e1b0", + "63769f20ce454f859db7991e5ee26a1e", + "4a2e7e51398b46f5b1edbe6c583b0fee", + "f4d67f931ad3433897fadfa57d893311", + "73199a8803954d4c8f838fc48005ba89", + "3e1536954e9d4902a82a8df5c7a15ce2", + "c67c69acb6c74d6c8d0e3606cc6c5771", + "262a15c6398b4c13a6ff6b39c38f30de", + "0e7ca1e4190c4561ba8d1123023a10c9", + "30d4eb4f89d24825ba8bdb11c083c20f", + "e1ba725e5a924bc38dde0b304f68a290", + "1e5d173e008c4d8ea8e0375d9080cf7d", + "1beb336970074f3cacd897cfe56bb48f", + "3dd37a0d116b42809a3d3c97a9937408", + "46bc292c79de49e39582bd380b7cb24d", + "d56efd640e23428eb12271af3bcce0ff", + "9c666db21edc429ca402c5571697b5f6", + "4959b758cb304a84817bcd67838f55c5", + "f49ac93865c54972a31484fa11585aae", + "819a7b4d4a2d449297e93555ddd66301", + "019363eaeada42b59a2da39e082741e2", + "e495962ffa0b4548b64720a4bc51f838", + "c61a2d13dbd744109e025a7d5a24f1cc", + "27fb8fc28cd84e809830d2f8ae6be768", + "8ccd3bc5d689460da73a88f9501d924a", + "9a59dbcb79754437aed83b9ea1d639b9", + "52201754b0bc4091a0c04876a5f99a07", + "810aeb2be35e453da30605bcc4c69de4", + "9ee433f566024fb69732263a5bfdc1fb", + "be5df40a08e4417798473164d0de1838", + "d2b91ac293db43e98589924b004eb8c3", + "da3e2830ceb14506aab3659f1b763394", + "574c168e6eef4fefa6e1c50c00fb5898", + "25ee11ebb92d4dd9a3c99335719a1372", + "a0a25b11371b4578af152f0c0ebda4fe", + "59b478cd6e1e4b4cbcdc7e086d4403a1", + "359a6b6e73ad4f00a12b26babf3556f5", + "ca34502cd5844bb1ba71e69fa1531657", + "8166ac4b22584cf399fdb1f198a48d1d", + "4151f29ec0e34c338289db98596b022f", + "f318b8c1e9424a15a601a9ce3ea12c38", + "f3177b640bf345aab255c08c8c808fdf", + "2b024e8e1f1741d5808d13dcf431a61e", + "822befe58307424f8d9b07f98565d561", + "36857dee08bf46678c092bf6f1de1468", + "33946ac898f84e0e9be4f3ec398b1720", + "d3d84f255b1e45ac9b2d9cba4810993f", + "a8b206b878de4a2bb7af14bd282e2987", + "63179fa57c984dce9b6dac7524c19447", + "f069a144805649f6b7358d855dccb7dd", + "864b4d9cb19a4910abfb448689a55603", + "700cbeb0644d40cc8954b93fea69bcbf", + "d523f309efad4cfaa156f7f670bac10f", + "729e27a0233b4f92b13d72df1b21b2fd", + "c1130e6bb0c5407aa9d75ae046978be3", + "845828ad19654d219b93303c33db8b0e", + "30699b746bd440b6b2dca85173c304d9", + "565bdcc9c19043a88bde9b0f9fc910f2", + "a984b3e68ad4400fa386e728f497beab", + "fd2d7856006340ad94d4ab083f231279", + "af4f189380ba4f118c4976d5379164b2", + "cc5bd76b7d2b485b8cc4cca9789f8673", + "b924d99d60bd4d3fb8697556ec41fcfb", + "9745b8f9bd1547ea8002b4fb0173c722", + "e41c1f571c4e4036ba6491e564c4a763", + "c49d6c830dda4284b13d1d99aef88a58", + "d9b0bc02f27e47948ddadf67797434ca", + "36d8576d8ec84d78a73249b9b111fa77", + "5ff63896f6d64d4a80a858583389d0c6", + "bbbb033fe303485a8898a739a2ac35e5", + "469897e960394259aaa7fb39f572dfae", + "71522dccde854cb0bb73fa03d1ceb453", + "7d50d57bfd5448aea2d741e07e50aae4", + "d99029215cbd465586da796ed65117e6", + "93ad654c0057407daeabc7b17ddae867", + "38b90c0e78734bf18cd9bf0dbfe5be7a", + "5ff8b298d4674a4eb6dbaaeef4322c56", + "72feff0edadf43feaf280db5bebe7ab9", + "d2d4857de45b43bbb0559a5303c34932", + "674a9a58dce84f5fbc2e8dcf65688775", + "0b474fbed20747f0ba7839fd69957095", + "b264371e632e4886889f4ed9c8b92560", + "636caf57d5b34d66b2a9175585e052f6", + "a3bad7d0f6f74c79adc956e392c1f5c1", + "438a1b373cbb47978490ede44c7a1da4", + "1192002b51a84875a22143b484e882a9", + "6466e87253f746a3bf927a70100453f3", + "c7acd7f863ea4b01999ba26691f6929f", + "628cd4e9b98b495abf3a5f6214d8ba71", + "094e9a7f211e4371ab0ccb98f0f03846", + "43d9c172daa742d7a055e5af64e1890c", + "7e92cbe0211d497b9657f715ab7636cc", + "46c83b4371164bd381052cf494922f69", + "43f2b15c83044b0a8799f4d21be491b4", + "3df6cc80de344feda6393649f943b715", + "039196b481e040b1a5536dc5219b38b0", + "b4fff616bb62448e8be771f4727f2d4f", + "2ffa0b0732f84680a8b1ef607b719b87", + "f81e6ca8ea364070a88ff187187d7522", + "377ca9656bb54bb2b2f43b1290854396", + "6ec9691a69b240c0842450b17a3dc1ed", + "cb48ecc9c9d64749b93eab50553a643b", + "b9267d3d61344b0e8796de908733f912", + "08daef067520467ebaa027cedc8a895d", + "85edf75bd18c4381842ae6f3fee8512c", + "9fd9a19e00884434a3ce89863109b5eb", + "9232db089f834931b8acf50bd5dc6cbc", + "616906ac2b2c4540a1b712d8f76da854", + "2ee5dd0c00584d389c4a64e633eb59c0", + "f4e3619de3eb49fc812a2dd52df2b59e", + "ac62107beb59424d911992fbb050a1bc", + "3aceddff8236439ab3e4db30b918b174", + "9406c36e359942e98b83e60fbe408d21", + "5836ddba34a041bbbfe66aa5d20568d0", + "29752439eb5940f7915d35712ad5f54e", + "8617aaae738348a0b5c0e5e4d334fd4e", + "75ecfa2ebd62408a8ff62871d82f47bf", + "b22ffc37c2e54886a38a31b28027b51f", + "5c7efa43d8f0469d929ae34f5425d3ce", + "85b6f496f3ee4628bdda3e5ff90a626e", + "6fd9f07536004cf583e9b45d6a4254ca", + "49e8bec8a16b4afb908bdcb0a1cae9e8", + "1125cd031b7e48a7bf1574dd00cb536d", + "790aa8865a434c548004a3c75729ac2e", + "d1b203b5a91a4e30a44844014606cf49", + "b01fed9deffd499e81f858a3a6271b0d", + "d8056d29ca3a49a78bedd271ccd62877", + "8c83d20bf2d64190b8333b90bbd30e00", + "8144b3ee3e06484db0139c7f1dd8a127", + "1583db27212246e78edb8f1734a29bcd", + "0165d1185a3c48f19a691a3446a80550", + "072bdd89f1044a1bbf0be4df161d0e3f", + "f063d9877ba14768ab3b0903454d75a5", + "75bb5a928b4848bfa1e7372afb3d2b6e", + "d2de8b6317a844d8b9040b62a986106a", + "bf61b54f22b14a7893fc3e8ab8286068", + "211318b770754f1ca00fcd98f0c2a708", + "0033cc4de608444a85bbfa6ac98db926", + "8e7120d574de4f6ab942881c510fd0f9", + "dac762b385f34159ae5418607b1b99f2", + "ca3227078b474e55bd06eb8733c2a148", + "ee8f105b1d7742479663eceb2ad7b909", + "7a736584b9704b36a849ad39046805df", + "ce44c8896ec54825aa19f1068c82f5a8", + "286eef7338c04b468a5aa0d7947e2dad", + "50818e80e8664f75b81e2626d1732c8f", + "50e2cd0062554ac3ade86556a4794fc4", + "e92614519ab04c2dac1da67c09a8234a", + "aa31e529f957420287dfbc8908b817de", + "04a054f8353c42f0845994baf5f8c7ae", + "72770cec3112462c8c4553eab5f797b7", + "373e437bbf374e03932d31a4f923ae71", + "b5b947f7ee8d46bcb9a378ce427cfe30", + "5e1bb37a7ab04844a110de9afec78d38", + "074d4db7152b4eee865b0da7931424f8", + "ad84eecca5f34785bedc78549fc89f1e", + "c3b6d7c4559c4605b577c557d8fd19a6", + "0891771ea54f48a6b2a825cb26f72248", + "7a2b2a28d8864dc4a3e16a4fbb589410", + "4d9e15cc91874a2a9ba4a2b3ec45523f", + "693a1327de424fcd8022509a49da4304", + "10fa9a7512304a83b08cbb36eb46e520", + "61ea423cfe2f48bb8b4478c099626a60", + "daa71dfede8a460fad2cabf486091e81", + "3b3d7ff2f32845498d98dbc70961031a", + "103ae87c16ac40caa5f881629e72f04a", + "f8d861f487124aee9685b522f43c0e64", + "08413cf2615c4fd38294ab7276559f0e", + "93231f2c440940108ba9dc87dce1c754", + "1d13021c96e74c6c907707d3e181ec86", + "f0232892bc4e41be991ce54034cc7abc", + "c454441dd0794a1e8756bdd72b961519", + "9eaf4e8079bd4305b4a695db020fdf12", + "17fd0089d6164d81aa733be84398e54d", + "11433cf089f047019243f73145aa3582", + "a8a86036f08241a1b437da0c2b024243", + "4aba135a18b64b9ca22bde3658b8a662", + "a6495b1a1ad04ce1afb1e2f469f0e561", + "dff371d78c5846b98e4cd353bdd7d68e", + "60091d1c1dde40bdbd66388528d50db8", + "74ff6a4b1ae241438e004c1e6c61b383", + "6357afa15d4b45a0884530e83b33d7da", + "fe0d69d63a0c48f185e90ced28939371", + "1846f4cecaa34ee1a3050fdca05c19d2", + "63c66aa8811e4940bd0e590f697ee0d4", + "fcd8145f440a48c9bd9c0400465ab07d", + "902856e8f30043a3b73722ff00cd1290", + "30d03fda9db6496ba3ceb9679c4931ba", + "79e9d9a4f5b0418eb3e2aa837d85780f", + "567610dc9de94bc19165af68da5eb265", + "db6a92356f844e348415d582d1ffb6d6", + "54c7e48a210a4785925bb117c82fb66f", + "bf00ba2c72d44ed9afad6ed1e04ba1ce", + "c70ed24c6f104e8cb5f52ecb937a97dc", + "3c2a0dc765d643c2b1b15cfe9346fe6f", + "bef2574ab0514c509bea09170b221eb7", + "14ed2bebe6804f8a9727c3dbc96a30aa", + "e4fadcc9836642f5aaeda766d6aa5197", + "5b8c462f3c2e4de1b14ed4de3a54ed49", + "edcd0ba8c8424186a164bfb276095cf5", + "edf53a46faec401da1f2d06a03f979f3", + "fb24d2f21ef342a690f34fa39857772e", + "4a2a273978074e17a8bcc365188a75ef", + "5e54a659200d472e98b2190e965aee45", + "b4b51345cd714db193598fe413b69ca2", + "2f8865a29dd54cc59ac96ec2a7e89612", + "f2e34d77b2644dbaacf98518dff660df", + "1e1f1a51e36445ee8305def8f4d26733", + "d6c61329d0524f99b707e7ddb0fb3963", + "6e2c9affc0ce40aab66cc6fe4cc9029e", + "329aa43ebb354435bae90ede556614b4", + "cb2d4a153bc24592a10a5beb7dcb567e", + "acf1419005004ec8b7a3ab0cb5b5be20", + "625953dfc8bf4ed5884701e5795ac134", + "e7697815ca02401fb448676ead592633", + "a1e7f48a470c4e23a285870c745395a8", + "ccc0c170d40d4acba18b5a2e34ec6130", + "89f860310d1144b4a3a6aad99a2388bf", + "454f9e86feab4312af6bdc456522ea14", + "fc1d80decdbe43239fcb100d4ae29987", + "da6904ae4c0d48c9804a8d34eeac2fa7", + "7876fdc2d58b464a8832f4123cd0ac84", + "4286ea80ab9b44d3bd74be460223cccd", + "80b99cdefc0c4a19a03e5b60977c5252", + "da5a5587ec6e4c9a81ea978502e5311b", + "bd7419fdcc404a469731b9761a51b7c2", + "56317622da284301b15654c89fb0496a", + "26aa2d4e43e24287bfa9e457d68132ab", + "3b988dc9683b4f11a8770fbdff99b4f7", + "2a71ebeaf0784fbeb052a1468ff0ca4d", + "3b5b15ec158a4f4c961399321dce9234", + "1b57bc90e7f342e08b93d6704cd5971f", + "8328e7bf86e34d60b480662c69cf4fbd", + "1de42c1e24334732a15d4e35bfb5b3ad", + "e42e9008c9f8419789c3da3728f8a8d0", + "c84f99822c8f4dab84cceea70c2875d9", + "57b5d6997ee742afb9f95f15108b105f", + "657cb831767d4e0ba12b5618c6873a53", + "90e56cfef4c64a79b67d8e280644a4c4", + "d6b0c8873fed4355a3c3f8e1b1aad2dc", + "37043db10f0848a09303744b39711a6e", + "b6bd0e5ef0964405bf110326f8b4dd76", + "045c33daa1aa4c0d99988fc8d30ddb1b", + "0ec8d78fc4744e038766a6fd26ee86f3", + "55ce376b87bc40688abd725562790b5a", + "0379f30bd3954b9fbb9153adfeb2668f", + "9206620cd48341a7b7c089462718736f", + "b512e866fba74546aa8483cc9b5a1cf4", + "06e6bd6db4ec4fa597b8b4af7ac727cb", + "5c7c9cbedf85418ca94ea5c2e486a587", + "fb49049931414cf3892c9b1ccf0c2bea", + "8ea31eff9315444d9fcef28853d91127", + "69ac9ce2a2ce4dd9bbef0c1baff0d211", + "45419da06ff94bc898956c50e786570b", + "c48f9e4435954f64a25fef8ad1a9f459", + "9f7a5df5e5ec4151b63bc8f48891b46c", + "b549783acd574ad1ae2006d61b810375", + "cadcb886bb42450d9359c25b5c7c75bd", + "2ac9ee865773429ca80db057e7337fbd", + "6243ea4cd1db43edacae8a3f257abf8d", + "a2709220d060416eaa11f23bb7475729", + "6603c5d36b8e403696bd78c01f6e81ed", + "33428b2324c94d66890d5abad6340086", + "5dc36030f0154b36913b0570422aee6d", + "2468442fd53445ce8143523b979cbb2e", + "ab319482a6634fa0b6b5c29b08fd067f", + "002df8624a9948dfa28234d313056317", + "45f90258a1b54457a5ac3d71de942faf", + "c5dc4d2c4f104d9f86cc640b6d9cdaca", + "758249fa6bb74a008eb604c4dac514d5", + "efe454c14cc34b4691bd247e8a2a4b91", + "18326325bc15408b866905bdd876209b", + "9c4bfc9f5932407293993cf19572e52c", + "f957ba6a1f0f4035a41e2018263668a2", + "0f62f523bc2f4919b9e4d4f324513b55", + "c2ef33d1fae24ef489b39591004d46d1", + "f6f5e0ca1b074a66a3d6987e9b740788" + ] + }, + "id": "4jD4c_LWW_Ef", + "outputId": "d5bd9ed6-61ba-408d-9a78-b97d6550e79e" + }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Comparing language usage with and without LoRA on 20 samples:\n", + "============================================================\n" + ] + }, + { + "data": { + "application/vnd.jupyter.widget-view+json": { + "model_id": "e7f13af32bfc4b06aa4efa5d81595639", + "version_major": 2, + "version_minor": 0 + }, + "text/plain": [ + "Processed prompts: 0%| | 0/1 [00:00 话不多说,直接开始! + +本文使用的测试环境为单张 A100,显存 80GB,可根据需求切换不同参数量的模型,实测4B 24G显存 is enough! +使用的框架为 Unsloth +![05-1](./images/05-1.png) +Unsloth 是一个极其强调资源节省的框架,把所有的资源节省做到了极致,具体来讲Unsloth能够将 Llama-3、Mistral、Phi-4 和 Gemma 等大型语言模型的微调速度提升 2 倍,内存占用减少 70%,并且准确率没有任何下降! +官方文档非常全面,详细指导了如何训练自己的定制模型。其中涵盖了安装和更新 Unsloth、创建数据集、运行和部署模型等基本要素。 Unsloth 让大家在本地或在 Google Colab 和 Kaggle 等平台上训练像 Llama 3 这样的模型变得极其简单。Unsloth简化了整个训练工作流程,包括模型加载、量化、训练、评估、运行、保存、导出,以及与 Ollama、llama.cpp 和 vLLM 等推理引擎的集成。 +Unsloth定期与 Hugging Face、Google 和 Meta 的团队合作,以修复 LLM 训练和模型中的错误。因此,当使用 Unsloth 进行训练或使用模型时,可以期待获得最准确的结果。 Unsloth 具有高度可定制性,允许更改聊天模板或数据集格式等内容。Unsloth还为视觉、文本转语音 (TTS)、BERT、强化学习 (RL) 等提供了预构建的脚本!此外,Unsloth支持所有训练方法和所有基于 Transformer 的模型。 + +## 教程概览 + +本教程将指导您完成 **DeepSeek-R1-Distill-Qwen3-8B 模型的 GRPO(Group Relative Policy Optimization)微调**,这是一种先进的强化学习技术,专门用于提升大语言模型在特定任务上的表现。 + +### 什么是GRPO? + +GRPO(Group Relative Policy Optimization)是一种强化学习优化技术,通过设计多个奖励函数来评估模型输出的不同方面,从而指导模型学习期望的行为模式。在数学推理任务中,GRPO可以帮助模型: + +- 学会按照特定格式输出答案 +- 提高推理过程的逻辑性 +- 增强答案的准确性 +- 改善输出的结构化程度 + +### 本教程的学习内容 + +1. **环境设置**: 安装Unsloth和相关依赖 +2. **模型加载**: 加载DeepSeek-R1-Distill-Qwen3-8B预训练模型 +3. **LoRA配置**: 设置高效的参数微调 +4. **数据处理**: 处理GSM8K数学推理数据集 +5. **格式设计**: 定义结构化的输出格式 +6. **奖励函数**: 设计多维度评估体系 +7. **GRPO训练**: 执行强化学习微调 +8. **效果验证**: 测试微调后的模型 +9. **模型保存**: 保存训练结果 +10. **可视化监控**: 使用SwanLab跟踪训练过程 + + + + +```python +# 安装依赖包 +# pip install unsloth vllm==0.8.5.post1 +``` + +```python +# 安装语言检测库 +# pip install langid -qq +``` + +```python +from unsloth import FastLanguageModel +import torch +max_seq_length = 1024 +lora_rank = 32 + +model, tokenizer = FastLanguageModel.from_pretrained( + model_name = "/opt/tiger/test0/DeepSeek-R1-0528-Qwen3-8B", + max_seq_length = max_seq_length, + load_in_4bit = True, # 对于LoRA 16位设置为False + fast_inference = True, # 启用vLLM快速推理 + max_lora_rank = lora_rank, + gpu_memory_utilization = 0.7, # 如果内存不足请减少此值 +) + +model = FastLanguageModel.get_peft_model( + model, + r = lora_rank, # 选择任何大于0的数字!建议8, 16, 32, 64, 128 + target_modules = [ + "q_proj", "k_proj", "v_proj", "o_proj", + "gate_proj", "up_proj", "down_proj", + ], + lora_alpha = lora_rank*2, # *2可以加速训练 + use_gradient_checkpointing = "unsloth", # 减少内存使用 + random_state = 3407, +) +``` + +### GRPO对话模板 + +```python +reasoning_start = None +reasoning_end = None +user_token = None +assistant_token = None + +for token in tokenizer.get_added_vocab().keys(): + if "think" in token and "/" in token: + reasoning_end = token + elif "think" in token: + reasoning_start = token + elif "user" in token: + user_token = token + elif "assistant" in token: + assistant_token = token + +system_prompt = \ +f"""你接到一个问题。 +请思考这个问题并提供你的解题过程。 +你必须用印尼语思考。""" +system_prompt +``` + +```python +print(tokenizer.apply_chat_template([ + {"role" : "user", "content" : "What is 1+1?"}, + {"role" : "assistant", "content" : f"I think it's 2.22"}, + {"role" : "user", "content" : "What is 1+1?"}, + {"role" : "assistant", "content" : f"I think it's 2.22"}, +], tokenize = False, add_generation_prompt = True)) +``` + +### 数据准备 +```python +from datasets import load_dataset +dataset = load_dataset("open-r1/DAPO-Math-17k-Processed", "en", split = "train") +dataset +``` + +让我们看看第一行数据: + +```python +dataset[0]["prompt"] +``` + +```python +dataset[0]["solution"] +``` + +在GSM8K中,我们注意到所有答案都有####标记,所以我们需要提取它。但对于Open R1数据集,我们可以跳过下面的处理。 + +```python +def extract_hash_answer(text): + # if "####" not in text: return None + # return text.split("####")[1].strip() + return text +extract_hash_answer(dataset[0]["solution"]) +``` + +让我们映射数据集!并查看第一行: + +```python +dataset = dataset.map(lambda x: { + "prompt" : [ + {"role": "system", "content": system_prompt}, + {"role": "user", "content": x["prompt"]}, + ], + "answer": extract_hash_answer(x["solution"]), +}) +dataset[0] +``` + +我们创建一个正则表达式格式来匹配推理部分和答案: + +```python +import re + +# 添加可选的EOS标记匹配 +solution_end_regex = rf"{reasoning_end}(.*)" + +match_format = re.compile(solution_end_regex, re.DOTALL) +match_format +``` + +我们验证它能正常工作: + +```python +match_format.findall( + "Let me think!"\ + f"Hence, the solution is 2.", +) +``` + +```python +match_format.findall( + "Let me think!"\ + f"\n\nHence, the solution is 2", +) +``` + +我们现在要创建一个奖励函数来完全匹配格式 - 如果成功匹配我们给3分: + +```python +def match_format_exactly(completions, **kwargs): + scores = [] + for completion in completions: + score = 0 + response = completion[0]["content"] + # 匹配是否完全符合格式! + if match_format.search(response) is not None: score += 3.0 + scores.append(score) + return scores +``` + +如果失败,我们希望在至少部分遵循格式时奖励模型,通过计算每个符号: + +```python +def match_format_approximately(completions, **kwargs): + scores = [] + for completion in completions: + score = 0 + response = completion[0]["content"] + # 计算看到多少个关键词 - 如果太多我们会惩罚! + # 如果我们看到1个,那么加一些分! + + # 不需要奖励因为我们总是预置它! + score += 0.5 if response.count(reasoning_start) == 1 else -1.0 + score += 0.5 if response.count(reasoning_end) == 1 else -1.0 + scores.append(score) + return scores +``` + +我们想要提取生成的答案,并奖励或惩罚它!我们还根据答案与真实答案的比率来奖励: + +```python +def check_answer(prompts, completions, answer, **kwargs): + question = prompts[0][-1]["content"] + responses = [completion[0]["content"] for completion in completions] + + extracted_responses = [ + guess.group(1) + if (guess := match_format.search(r)) is not None else None \ + for r in responses + ] + + scores = [] + for guess, true_answer in zip(extracted_responses, answer): + score = 0 + if guess is None: + scores.append(-2.0) + continue + # 正确答案得5分! + if guess == true_answer: + score += 5.0 + # 如果看到空格但奖励较少 + elif guess.strip() == true_answer.strip(): + score += 3.5 + else: + # 我们也通过比率奖励接近的答案! + # 即如果答案在某个范围内,奖励它! + try: + ratio = float(guess) / float(true_answer) + if ratio >= 0.9 and ratio <= 1.1: score += 2.0 + elif ratio >= 0.8 and ratio <= 1.2: score += 1.5 + else: score -= 2.5 # 惩罚错误答案 + except: + score -= 4.5 # 惩罚 + scores.append(score) + return scores +``` + +有时答案可能不是1个数字,而是像句子一样,例如"解决方案是$20" -> 我们提取20。 + +我们还移除可能的逗号,例如123,456 + +```python +match_numbers = re.compile( + r".*?[\s]{0,}([-]?[\d\.\,]{1,})", + flags = re.MULTILINE | re.DOTALL +) +print(match_numbers.findall(" 0.34 ")) +print(match_numbers.findall(" 123,456 ")) +print(match_numbers.findall(" -0.234 ")) +print(match_numbers.findall("17")) +``` + +最后,我们将尝试强制思考过程使用印尼语。这是DeepSeek R1论文中使用的`语言一致性奖励`的简单版本 + +```python +import langid + +def get_lang(text: str) -> str: + if not text: + return "und" + lang, _ = langid.classify(text) + return lang + + +print(get_lang("Hello, How are you")) # 这应该返回en +print(get_lang("Aku berpikir kalau aku adalah kamu")) # 这应该返回id +print(get_lang("我在这里")) # 这应该返回zh +``` + +```python +import re + +def format_and_language_reward_func(completions, **kwargs): + scores = [] + + for completion_item in completions: + if not completion_item or not isinstance(completion_item[0], dict) or "content" not in completion_item[0]: + scores.append(-5.0) + print(f"警告:格式错误的完成项,分配默认低分: {completion_item}") + continue + + content = completion_item[0]["content"] + + lang = get_lang(content) + + if lang == 'id': + score = 5.0 + elif lang == 'en': + score = -3.0 + elif lang == 'zh': + score = -3.0 + else: + score = -5.0 + + scores.append(score) + + return scores +``` + +```python +prompts = [ + [{"role": "assistant", "content": "What is the result of (1 + 2) * 4?"}], + [{"role": "assistant", "content": "What is the result of (3 + 1) * 2?"}], +] +completions = [ + [{"role": "assistant", "content": "The sum of 1 and 2 is 3, which we multiply by 4 to get 12.(1 + 2) * 4 = 12"}], + [{"role": "assistant", "content": "The sum of 3 and 1 is 4, which we multiply by 2 to get 8. So (3 + 1) * 2 = 8."}], +] +format_and_language_reward_func(prompts=prompts, completions=completions) +``` + +我们现在准备主函数,它将打印生成的响应和真实答案,以及另一个奖励函数,通过`float`将文本转换为浮点数并查看是否相同。 + +```python +global PRINTED_TIMES +PRINTED_TIMES = 0 +global PRINT_EVERY_STEPS +PRINT_EVERY_STEPS = 5 + +def check_numbers(prompts, completions, answer, **kwargs): + question = prompts[0][-1]["content"] + responses = [completion[0]["content"] for completion in completions] + + extracted_responses = [ + guess.group(1) + if (guess := match_numbers.search(r)) is not None else None \ + for r in responses + ] + + scores = [] + # 只在每几步打印一次 + global PRINTED_TIMES + global PRINT_EVERY_STEPS + if PRINTED_TIMES % PRINT_EVERY_STEPS == 0: + print( + '*'*20 + f"问题:\n{question}", f"\n答案:\n{answer[0]}", f"\n响应:\n{responses[0]}", f"\n提取的:\n{extracted_responses[0]}" + ) + PRINTED_TIMES += 1 + + for guess, true_answer in zip(extracted_responses, answer): + if guess is None: + scores.append(-2.5) + continue + # 转换为数字 + try: + true_answer = float(true_answer.strip()) + # 移除逗号,如123,456 + guess = float(guess.strip().replace(",", "")) + scores.append(3.5 if guess == true_answer else -1.5) + except: + scores.append(0) + continue + return scores +``` + +获取前90%的提示长度,这样我们就不会意外截断它们! + +即我们将移除前10%的长提示。 + +```python +tokenized = dataset.map( + lambda x: {"tokens" : tokenizer.apply_chat_template(x["prompt"], add_generation_prompt = True, tokenize = True)}, + batched = True, +) +print(tokenizer.decode(tokenized[0]["tokens"])) +tokenized = tokenized.map(lambda x: {"L" : len(x["tokens"])}) + +import numpy as np +maximum_length = int(np.quantile(tokenized["L"], 0.9)) +print("最大长度 = ", maximum_length) + +# 只过滤小于90%最大长度的样本 +dataset = dataset.select(np.where(np.array(tokenized["L"]) <= maximum_length)[0]) +del tokenized +``` + + +### 训练模型 + +现在设置GRPO训练器和所有配置! + +```python +max_prompt_length = maximum_length + 1 # +1以防万一! +max_completion_length = max_seq_length - max_prompt_length + +from vllm import SamplingParams +vllm_sampling_params = SamplingParams( + min_p = 0.1, + top_p = 1.0, + top_k = -1, + seed = 3407, + stop = [tokenizer.eos_token], + include_stop_str_in_output = True, +) + +from trl import GRPOConfig, GRPOTrainer +training_args = GRPOConfig( + vllm_sampling_params = vllm_sampling_params, + temperature = 1.0, + learning_rate = 5e-6, + weight_decay = 0.01, + warmup_ratio = 0.1, + lr_scheduler_type = "linear", + optim = "adamw_8bit", + logging_steps = 1, + per_device_train_batch_size = 1, + gradient_accumulation_steps = 1, # 增加到4以获得更平滑的训练 + num_generations = 4, # 如果内存不足请减少 + max_prompt_length = max_prompt_length, + max_completion_length = max_completion_length, + # num_train_epochs = 1, # 对于完整训练运行设置为1 + max_steps = 100, + save_steps = 100, + report_to = "swanlab", # 可以使用Weights & Biases + output_dir = "outputs", + + # 用于可选的训练+评估 + # fp16_full_eval = True, + # per_device_eval_batch_size = 4, + # eval_accumulation_steps = 1, + # eval_strategy = "steps", + # eval_steps = 1, +) +``` + +让我们运行训练器!如果你向上滚动,你会看到一个奖励表格。目标是看到`reward`列增加! + +你可能需要等待150到200步才能看到任何效果。前100步你可能会得到0奖励。请耐心等待! + +| 步骤 | 训练损失 | 奖励 | 奖励标准差 | 完成长度 | kl | +|------|---------------|-----------|------------|-------------------|----------| +| 1 | 0.000000 | 0.125000 | 0.000000 | 200.000000 | 0.000000 | +| 2 | 0.000000 | 0.072375 | 0.248112 | 200.000000 | 0.000000 | +| 3 | 0.000000 | -0.079000 | 0.163776 | 182.500000 | 0.000005 | + +```python +# 用于可选的训练+评估 +# new_dataset = dataset.train_test_split(test_size = 0.01) + +trainer = GRPOTrainer( + model = model, + processing_class = tokenizer, + reward_funcs = [ + match_format_exactly, + match_format_approximately, + check_answer, + check_numbers, + format_and_language_reward_func, + ], + args = training_args, + train_dataset = dataset, + + # 用于可选的训练+评估 + # train_dataset = new_dataset["train"], + # eval_dataset = new_dataset["test"], +) +trainer.train() +``` + + +### 推理 +现在让我们试试刚刚训练的模型!首先,让我们先试试没有经过GRPO训练的模型: + +```python +text = "What is the sqrt of 101?" + +from vllm import SamplingParams +sampling_params = SamplingParams( + temperature = 1.0, + top_k = 50, + max_tokens = 1024, +) +output = model.fast_generate( + [text], + sampling_params = sampling_params, + lora_request = None, +)[0].outputs[0].text + +output +``` + +现在使用我们刚刚用GRPO训练的LoRA - 我们首先保存LoRA! + +```python +model.save_lora("grpo_lora") +``` + +验证LoRA确实被训练了! + +```python +from safetensors import safe_open + +tensors = {} +with safe_open("grpo_lora/adapter_model.safetensors", framework = "pt") as f: + # 验证A和B都非零 + for key in f.keys(): + tensor = f.get_tensor(key) + n_zeros = (tensor == 0).sum() / tensor.numel() + assert(n_zeros.item() != tensor.numel()) +``` + +现在我们加载LoRA并测试。我们在不使用自定义系统提示的情况下进行测试,这应该不会(或很少)影响模型的原始推理能力: + +```python +messages = [ + {"role": "user", "content": "Solve (x + 2)^2 = 0"}, +] + +text = tokenizer.apply_chat_template( + messages, + add_generation_prompt = True, # 生成时必须添加 + tokenize = False, +) +from vllm import SamplingParams +sampling_params = SamplingParams( + temperature = 1.0, + top_k = 50, + max_tokens = 2048, +) +output = model.fast_generate( + text, + sampling_params = sampling_params, + lora_request = model.load_lora("grpo_lora"), +)[0].outputs[0].text + +output +``` + +接下来,让我们使用系统提示进行测试,这应该使用新语言: + +```python +messages = [ + {"role": "system", "content": system_prompt}, + {"role": "user", "content": "Solve (x + 2)^2 = 0"}, +] + +text = tokenizer.apply_chat_template( + messages, + add_generation_prompt = True, # 生成时必须添加 + tokenize = False, +) +from vllm import SamplingParams +sampling_params = SamplingParams( + temperature = 1.0, + top_k = 50, + max_tokens = 2048, +) +output = model.fast_generate( + text, + sampling_params = sampling_params, + lora_request = model.load_lora("grpo_lora"), +)[0].outputs[0].text + +output +``` + +让我们比较使用系统提示但不使用LoRA的结果 + +```python +messages = [ + {"role": "system", "content": system_prompt}, + {"role": "user", "content": "Solve (x + 2)^2 = 0"}, +] + +text = tokenizer.apply_chat_template( + messages, + add_generation_prompt = True, # 生成时必须添加 + tokenize = False, +) +from vllm import SamplingParams +sampling_params = SamplingParams( + temperature = 1.0, + top_k = 50, + max_tokens = 2048, +) +output = model.fast_generate( + text, + sampling_params = sampling_params, + lora_request = None, +)[0].outputs[0].text + +output +``` + +让我们取20个样本,比较使用LoRA和不使用LoRA的情况,看看哪一个有更好的正确语言使用量 + +```python +sample_dataset = dataset.shuffle(seed = 3407).select(range(20)) +sample_dataset +``` + +```python +with_lora_id_count = 0 +without_lora_id_count = 0 + +print("在20个样本上比较使用和不使用LoRA的语言使用情况:") +print("=" * 60) + +for i, sample in enumerate(sample_dataset): + messages = [ + {"role": "system", "content": system_prompt}, + {"role": "user", "content": sample["prompt"][1]["content"]}, + ] + + text = tokenizer.apply_chat_template( + messages, + add_generation_prompt=True, + tokenize=False, + ) + + output_with_lora = model.fast_generate( + text, + sampling_params=sampling_params, + lora_request=model.load_lora("grpo_lora"), + )[0].outputs[0].text + + output_without_lora = model.fast_generate( + text, + sampling_params=sampling_params, + lora_request=None, + )[0].outputs[0].text + + lang_with_lora = get_lang(output_with_lora) + lang_without_lora = get_lang(output_without_lora) + + if lang_with_lora == 'id': + with_lora_id_count += 1 + if lang_without_lora == 'id': + without_lora_id_count += 1 + + # 每5个样本打印进度 + if (i + 1) % 5 == 0: + print(f"已处理 {i + 1}/20 个样本...") + +print("\n" + "=" * 60) +print("结果:") +print(f"使用LoRA - 印尼语响应: {with_lora_id_count}/20 ({with_lora_id_count/20*100:.1f}%)") +print(f"不使用LoRA - 印尼语响应: {without_lora_id_count}/20 ({without_lora_id_count/20*100:.1f}%)") +print(f"改进: 使用LoRA增加了{with_lora_id_count - without_lora_id_count}个印尼语响应") +``` + +我们的推理模型要好得多 - 它不总是正确的,因为我们只训练了大约一个小时 - 如果我们延长序列长度并训练更长时间会更好! + + +### 保存为float16格式用于VLLM + +我们还支持直接保存为`float16`。选择`merged_16bit`用于float16或`merged_4bit`用于int4。我们还允许`lora`适配器作为备选方案。使用`push_to_hub_merged`上传到你的Hugging Face账户!你可以去https://huggingface.co/settings/tokens获取你的个人令牌。 + +```python +# 合并为16位 +if False: model.save_pretrained_merged("model", tokenizer, save_method = "merged_16bit",) +if False: model.push_to_hub_merged("hf/model", tokenizer, save_method = "merged_16bit", token = "") + +# 合并为4位 +if False: model.save_pretrained_merged("model", tokenizer, save_method = "merged_4bit",) +if False: model.push_to_hub_merged("hf/model", tokenizer, save_method = "merged_4bit", token = "") + +# 仅LoRA适配器 +if False: + model.save_pretrained("model") + tokenizer.save_pretrained("model") +if False: + model.push_to_hub("hf/model", token = "") + tokenizer.push_to_hub("hf/model", token = "") +``` + +### GGUF / llama.cpp 转换 +要保存为`GGUF` / `llama.cpp`,我们现在原生支持它!我们克隆`llama.cpp`并默认保存为`q8_0`。我们允许所有方法,如`q4_k_m`。使用`save_pretrained_gguf`进行本地保存,使用`push_to_hub_gguf`上传到HF。 + +一些支持的量化方法(完整列表在我们的[Wiki页面](https://github.com/unslothai/unsloth/wiki#gguf-quantization-options)): +* `q8_0` - 快速转换。高资源使用,但通常可接受。 +* `q4_k_m` - 推荐。对一半的attention.wv和feed_forward.w2张量使用Q6_K,其他使用Q4_K。 +* `q5_k_m` - 推荐。对一半的attention.wv和feed_forward.w2张量使用Q6_K,其他使用Q5_K。 + +[**新功能**] 要微调并自动导出到Ollama,试试我们的[Ollama笔记本](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3_(8B)-Ollama.ipynb) + +```python +# 保存为8位Q8_0 +if False: model.save_pretrained_gguf("model", tokenizer,) +# 记住去https://huggingface.co/settings/tokens获取令牌! +# 并将hf改为你的用户名! +if False: model.push_to_hub_gguf("hf/model", tokenizer, token = "") + +# 保存为16位GGUF +if False: model.save_pretrained_gguf("model", tokenizer, quantization_method = "f16") +if False: model.push_to_hub_gguf("hf/model", tokenizer, quantization_method = "f16", token = "") + +# 保存为q4_k_m GGUF +if False: model.save_pretrained_gguf("model", tokenizer, quantization_method = "q4_k_m") +if False: model.push_to_hub_gguf("hf/model", tokenizer, quantization_method = "q4_k_m", token = "") + +# 保存为多个GGUF选项 - 如果你想要多个选项会快很多! +if False: + model.push_to_hub_gguf( + "hf/model", # 将hf改为你的用户名! + tokenizer, + quantization_method = ["q4_k_m", "q8_0", "q5_k_m",], + token = "", + ) +``` + + +### 本试验的试验记录 + +#### GRPO阶段 + +![05-3](./images/05-3.png) +400个step之后loss会有明显变化 + +## 教程总结 + +🎉 恭喜!你已经成功完成了DeepSeek-R1-Distill-Qwen3-8B的GRPO微调教程。 + +### 本教程涵盖的核心概念: + +1. **GRPO微调**: 使用奖励函数指导模型学习特定输出格式 +2. **LoRA技术**: 高效的参数微调方法,节省显存和时间 +3. **奖励函数设计**: 多层次评估体系,从格式到内容的全面评价 +4. **结构化输出**: 训练模型按照特定格式输出推理过程和答案 +5. **SwanLab监控**: 实时跟踪训练进度和指标变化 + +### 学到的技能: + +- ✅ 设置GRPO训练环境 +- ✅ 设计多维度奖励函数 +- ✅ 配置LoRA参数进行高效微调 +- ✅ 处理数学推理数据集 +- ✅ 监控和分析训练过程 +- ✅ 保存和部署微调模型 + +### 进一步探索: + +1. **调整奖励函数**: 设计更复杂的评估机制 +2. **扩展数据集**: 使用更大或不同类型的数据集 +3. **优化参数**: 尝试不同的LoRA配置和训练参数 +4. **模型评估**: 在测试集上系统评估模型性能 +5. **应用部署**: 将模型集成到实际应用中 + +### 注意事项: + +- 本教程使用了较少的训练步数作为演示,实际应用中建议使用更多步数 +- 可以根据显存情况调整批次大小和生成数量 +- SwanLab提供了丰富的可视化功能,建议深入探索 + +感谢你的学习!如果有任何问题,欢迎查看SwanLab的实验记录或重新运行代码。 + +# 总结 + +Congratulations!看到了这,你已经初步实现了一个简单的RL实战,掌握了使用 Unsloth 对 Gemma3 这类大模型进行 GRPO 微调的具体操作步骤,更能体会到 Unsloth 在大幅提升训练速度、显著降低显存占用方面的强大优势,从而使在有限资源下进行复杂强化学习实验成为可能!如果支持我们的工作希望得到你的star!!这是我们持续更新的最大动力!!! + +# 相关链接 + +- 完整可运行的代码:[Github](https://github.com/datawhalechina/self-llm/blob/master/models/DeepSeek-R1-Distill-Qwen/05-DeepSeek-R1-Distill-Qwen3-8B-GRPO及swanlab可视化.ipynb) +- 综述:https://arxiv.org/abs/2001.06921 +- deepseek-r1:https://arxiv.org/abs/2501.12948 +- 数学原理:https://blog.csdn.net/weixin\_38991876/article/details/146474767 +- Unsloth:https://docs.unsloth.ai/ diff --git a/models/DeepSeek-R1-Distill-Qwen/images/05-1.png b/models/DeepSeek-R1-Distill-Qwen/images/05-1.png new file mode 100644 index 0000000..ff0f766 Binary files /dev/null and b/models/DeepSeek-R1-Distill-Qwen/images/05-1.png differ diff --git a/models/DeepSeek-R1-Distill-Qwen/images/05-2.png b/models/DeepSeek-R1-Distill-Qwen/images/05-2.png new file mode 100644 index 0000000..27d999a Binary files /dev/null and b/models/DeepSeek-R1-Distill-Qwen/images/05-2.png differ diff --git a/models/DeepSeek-R1-Distill-Qwen/images/05-3.png b/models/DeepSeek-R1-Distill-Qwen/images/05-3.png new file mode 100644 index 0000000..cda93f6 Binary files /dev/null and b/models/DeepSeek-R1-Distill-Qwen/images/05-3.png differ