更新英文文档 Join Us (#1411)

This commit is contained in:
Henry Wang
2019-01-17 11:25:38 +08:00
committed by DIYgod
parent 05c04cbe06
commit 1a31db94c3
2 changed files with 437 additions and 120 deletions
+412 -106
View File
@@ -8,9 +8,416 @@ We welcome all pull requests. Suggestions and feedback are also welcomed [here](
## Submit new RSS source
1. Add a new route in [/lib/router.js](https://github.com/DIYgod/RSSHub/blob/master/lib/router.js)
### Step 1: Code the script
1. Add the script to the corresponding directory [/routes/](https://github.com/DIYgod/RSSHub/tree/master/routes)
Firstly, add a .js file for the new route in [/lib/router.js](https://github.com/DIYgod/RSSHub/blob/master/lib/router.js)
#### Acquiring Data
- Typically the data are acquired via HTTP requests (via API or webpage) sent by [axios](https://github.com/axios/axios)
- Occasionally [puppeteer](https://github.com/GoogleChrome/puppeteer) is required for browser stimulation and page rendering in order to acquire the data
- The acquired data are most likely in JSON or HTML format
- For HTML format, [cheerio](https://github.com/cheeriojs/cheerio) is used for further processing
- Below is a list of data acquisition methods, ordered by the **「level of recommendation」**
1. **Acquire data via API using axios**
Example:[/lib/routes/bilibili/coin.js](https://github.com/DIYgod/RSSHub/blob/master/lib/routes/bilibili/coin.js)。
Acquiring data via the official API provided by the data source using axios:
```js
// Initiate a HTTP GET request
const response = await axios({
method: 'get',
url: `https://api.bilibili.com/x/space/coin/video?vmid=${uid}&jsonp=jsonp`,
});
const data = response.data.data; // response.data is the data object returned from the previous request
// The object contains a nested object called data, thus response.data.data is the actual data needed here
```
One of the leaf objects (response.data.data[0]):
```json
{
"aid": 33614333,
"videos": 2,
"tid": 20,
"tname": "宅舞",
"copyright": 1,
"pic": "http://i0.hdslb.com/bfs/archive/5649d7fe6ff7f7b431300fc1a0db80d3f174cacd.jpg",
"title": "【赤九玖】响喜乱舞【和我一起狂舞吧,团长大人(✧◡✧)】",
"pubdate": 1539259203,
"ctime": 1539249536,
"desc": "编舞出处:av31984673\n真心好喜欢这个舞和这首歌,居然恰巧被邀请跳了,感谢《苍之纪元》官方的邀请。这次cos的是游戏的新角色缪斯。然而时间有限很多地方还有很多不足。也没跳够,以后私下还会继续练习,希望能学到更多动作,也能为了有机会把它跳的更好。 \n摄影:绯山圣瞳九命猫 \n后期:炉火"
// some more data....
}
```
Processing the data further to generate objects in accordance with RSS specification, mainly title, link, description, publish time, then assign them to ctx.state.data, [produce RSS feed](#produce-rss-feed):
```js
ctx.state.data = {
// the source title
title: `${name} 的 bilibili 投币视频`,
// the source link
link: `https://space.bilibili.com/${uid}`,
// the source description
description: `${name} 的 bilibili 投币视频`,
// iterate through all leaf objects
item: data.map((item) => ({
// the article title
title: item.title,
// the article content
description: `${item.desc}<br><img referrerpolicy="no-referrer" src="${item.pic}">`,
// the article publish time
pubDate: new Date(item.time * 1000).toUTCString(),
// the article link
link: `https://www.bilibili.com/video/av${item.aid}`,
})),
};
// the route is now done
```
2. **Acquire data via HTML webpage using axios**
Data have to be acquired via HTML webpage if **no API was provided**, for example: [/lib/routes/jianshu/home.js](https://github.com/DIYgod/RSSHub/blob/master/lib/routes/jianshu/home.js)。
Acquiring data by scrapping the HTML using axios:
```js
// Initiate a HTTP GET request
const response = await axios({
method: 'get',
url: 'https://www.jianshu.com',
});
const data = response.data; // response.data is the entire HTML source of the target page, returned from the previous request
```
Parsing the HTML using cheerio:
```js
const $ = cheerio.load(data); // Load the HTML returned into cheerio
const list = $('.note-list li').get();
// use cheerio selector, select all 'li' elements with 'class="note-list"', the result is an array of cheerio node objects
// use cheerio get() method to transform a cheerio node object array into a node array
// PS:every cheerio node is a HTML DOM
// PPS:cheerio selector is almost identical to jquery selector
// Refer to cheerio docs:https://cheerio.js.org/
```
Use /jianshu/utils.js class to extract full-text:
```js
const result = await util.ProcessFeed(list, ctx.cache);
```
The logic for full-text extraction in /jianshu/utils.js class:
```js
// define a function to load the article content
async function load(link) {
// get the article asynchronously
const response = await axios.get(link);
// load the article content
const $ = cheerio.load(response.data);
// parse the date
const date = new Date(
$('.publish-time')
.text()
.match(/\d{4}.\d{2}.\d{2} \d{2}:\d{2}/)
);
// handle the timezone
const timeZone = 8;
const serverOffset = date.getTimezoneOffset() / 60;
const pubDate = new Date(date.getTime() - 60 * 60 * 1000 * (timeZone + serverOffset)).toUTCString();
// extract the full-text
const description = $('.show-content-free').html();
// return the parsed result
return { description, pubDate };
}
// use Promise.all() to initiate requests in parallel
const result = await Promise.all(
// loop through every article
list.map(async (item) => {
const $ = cheerio.load(item);
const $title = $('.title');
// resolve the absolute URL
const itemUrl = url.resolve(host, $title.attr('href'));
// form a new object to hold the data
const single = {
title: $title.text(),
link: itemUrl,
author: $('.nickname').text(),
guid: itemUrl,
};
// use tryGet() to query the cache
// if the query returns no result, query the data source via load() to get article content
const other = await caches.tryGet(itemUrl, async () => await load(itemUrl), 3 * 60 * 60);
// merge two objects to form the final output
return Promise.resolve(Object.assign({}, single, other));
})
);
```
Assign the value of `result` to `ctx.state.data`
```js
ctx.state.data = {
title: '简书首页',
link: 'https://www.jianshu.com',
// select "content" property of <meta name="description">
description: $('meta[name="description"]').attr('content'),
item: result,
};
// the route is now done
```
3. **Acquire data via page rendering using puppeteer**
::: tip tips
This method consumes more resources and is less performant, use only when the above methods failed to acquire data, otherwise your pull requests will be rejected!
:::
Seldomly, data source **provides no API and the page requires rendering** to acquire data, for example: [/lib/routes/sspai/series.js](https://github.com/DIYgod/RSSHub/blob/master/lib/routes/sspai/series.js)
```js
// use puppeteer util class, initialise a browser instance
const browser = await require('../../utils/puppeteer')();
// open a new page
const page = await browser.newPage();
// access the target link
const link = 'https://sspai.com/series';
await page.goto(link);
// render the page
const html = await page.evaluate(
() =>
// process on the rendered page
document.querySelector('div.new-series-wrapper').innerHTML
);
// shutdown the browser
browser.close();
```
Parsing the HTML using cheerio:
```js
const $ = cheerio.load(html); // Load the HTML returned into cheerio
const list = $('div.item'); // // use cheerio selector, select all 'div class="item"' elements, the result is an array of cheerio node objects
```
Assign the value to `ctx.state.data`
```js
ctx.state.data = {
title: '少数派 -- 最新上架付费专栏',
link,
description: '少数派 -- 最新上架付费专栏',
item: list
.map((i, item) => ({
// the article title
title: $(item)
.find('.item-title a')
.text()
.trim(),
// the article link
link: url.resolve(
link,
$(item)
.find('.item-title a')
.attr('href')
),
// the article author
author: $(item)
.find('.item-author')
.text()
.trim(),
}))
.get(), // use cheerio get() method to transform a cheerio node object array into a node array
};
// the route is now done
// PS: the route acts as a notifier of new articles, it does not provide access to the content behind the paywall, thus not content were fetched
```
---
#### Enable Caching
By default there is a global caching period set in `lib/config.js`, some sources might have a low update frequency, a longer caching period should be set.
- Save to cache:
```js
ctx.cache.set((key: string), (value: string), (time: number)); // time is the caching period in seconds
```
- Access the cache:
```js
const value = await ctx.cache.get((key: string));
```
For example: [/lib/routes/zhihu/daily.js](https://github.com/DIYgod/RSSHub/blob/master/lib/routes/zhihu/daily.js), the full-text extraction will be triggered even when the article was not updated.
Given the update frequency is known, set the appropriate caching period to reuse the cache, will save resources and improve performance.
```js
const key = 'daily' + story.id; // story.id is the unique identifier of each article
ctx.cache.set(key, item.description, 24 * 60 * 60); // set the caching period to 24 hours * 60 minutes * 60 seconds = 86,400 seconds = 1 day
```
When the identical requests come in, reuse the cache:
```js
const key = 'daily' + story.id;
const value = await ctx.cache.get(key); // query the cache to find the unique identifier
if (value) {
// return the cached data
item.description = value; // assign the cached data
} else {
// no cached found
// initiate request to the data source
}
```
---
#### Produce RSS Feed
Assign the acquired data to ctx.state.data, the middleware [template.js](https://github.com/DIYgod/RSSHub/blob/master/middleware/template.js) will then process the data and render the RSS output [/views/rss.art](https://github.com/DIYgod/RSSHub/blob/master/views/rss.art), the list of parameters:
```js
ctx.state.data = {
title: '', // The feed title
link: '', // The feed link
description: '', // The feed description
language: '', // The language of the channel
item: [
// An article of the feed
{
title: '', // The article title
author: '', // Author of the article
category: '', // Article category
// category: [''], // Multiple category
description: '', // The article summary or content
pubDate: '', // The article publishing datetime
guid: '', // The article unique identifier, optional, default to the article link below
link: '', // The article link
},
],
};
```
##### Podcast feed
Used for audio feed, these **additional** data are in accordance with many podcast players' subscription format:
```js
ctx.state.data = {
itunes_author: '', // The channel's author, you must fill this data.
itunes_category: '', // Channel category
image: '', // Channel's image
item: [
{
itunes_item_image: '', // The item image
enclosure_url: '', // The item's audio link
enclosure_length: '', // The audio length in seconds.
enclosure_type: '', // Common types are: 'audio/mpeg' for .mp3, 'audio/x-m4a' for .m4a 'video/mp4' for .mp4
},
],
};
```
##### BT/Magnet feed
Used for downloader feed, these **additional** data are in accordance with many downloaders' subscription format to trigger automated download:
```js
ctx.state.data = {
item: [
{
enclosure_url: '', // Magnet URI
enclosure_length: '', // The audio length, the unit is seconds, optional
enclosure_type: 'application/x-bittorrent', // Fixed to 'application/x-bittorrent'
},
],
};
```
</details>
---
### Step 2: Add the script into router
Add the script into [/lib/router.js](https://github.com/DIYgod/RSSHub/blob/master/lib/router.js)
#### Example
1. [bilibili/bangumi](https://github.com/DIYgod/RSSHub/blob/master/lib/routes/bilibili/bangumi.js)
| Name | Description |
| ---------------------------------- | ---------------------------------------------------------------------------------- |
| Route | `/bilibili/bangumi/:seasonid` |
| Data Source | bilibili |
| Route Name | bangumi |
| Parameter 1 | :seasonid required |
| Parameter 2 | n/a |
| Parameter 3 | n/a |
| Route Path | `./routes/bilibili/bangumi` |
| the complete code in lib/router.js | `router.get('/bilibili/bangumi/:seasonid', require('./routes/bilibili/bangumi'));` |
2. [github/issue](https://github.com/DIYgod/RSSHub/blob/master/lib/routes/github/issue.js)
| Name | Description |
| ---------------------------------- | ---------------------------------------------------------------------------- |
| Route | `/github/issue/:user/:repo` |
| Data Source | github |
| Route Name | issue |
| Parameter 1 | :user, required |
| Parameter 2 | :repo, required |
| Parameter 3 | n/a |
| Route Path | `./routes/github/issue` |
| the complete code in lib/router.js | `router.get('/github/issue/:user/:repo', require('./routes/github/issue'));` |
3. [embassy](https://github.com/DIYgod/RSSHub/blob/master/lib/routes/embassy/index.js)
| Name | Description |
| ---------------------------------- | ---------------------------------------------------------------------------- |
| Route | `/embassy/:country/:city?` |
| Data Source | embassy |
| Route Name | n/a |
| Parameter 1 | :country, required |
| Parameter 2 | ?city, optional |
| Parameter 3 | n/a |
| Route Path | `./routes/embassy/index` |
| the complete code in lib/router.js | `router.get('/embassy/:country/:city?', require('./routes/embassy/index'));` |
---
### Step 3: Add the documentation
1. Update [Documentation (/docs/en/README.md) ](https://github.com/DIYgod/RSSHub/blob/master/docs/en/README.md), preview the docs via `npm run docs:dev`
@@ -61,112 +468,11 @@ We welcome all pull requests. Suggestions and feedback are also welcomed [here](
</routeEn>
1) Execute `npm run format` to lint the code before you commit and open a pull request
1. Execute `npm run format` to lint the code before you commit and open a pull request
## Write the script
RSSHub provides 3 methods for acquiring data, these methods are sorted by **recommended**:
### Access the target data source API
Use [axios](https://github.com/axios/axios) to access the target data source API, assign the acquired title, link, description and datetime to ctx.state.data (refer to Data for the list of parameters) , typically it looks like this: [/routes/bilibili/bangumi.js](https://github.com/DIYgod/RSSHub/blob/master/routes/bilibili/bangumi.js)
### Acquire data from HTML
If an API is not provided, data need to be scraped from HTML. Use [axios](https://github.com/axios/axios) to acquire the HTML and then use [cheerio](https://github.com/cheeriojs/cheerio) for scraping the relevant data and assign them to ctx.state.data, typically it looks like this: [/routes/jianshu/home.js](https://github.com/DIYgod/RSSHub/blob/master/routes/jianshu/home.js)
### Page rendering
::: tip tip
This method is comparatively less performant and consumes more resources, only use when necessary or your pull requests might be rejected.
:::
Some websites provides no API and pages require rendering too, use [puppeteer](https://github.com/GoogleChrome/puppeteer) render the pages via Headless Chrome and then use [cheerio](https://github.com/cheeriojs/cheerio) for scraping the relevant data and assign them to ctx.state.data, typically it looks like this: [/routes/sspai/series.js](https://github.com/DIYgod/RSSHub/blob/master/routes/sspai/series.js)
### Enable caching
All routes has a default cache expiry time set in `lib/config.js`, it should be increased when the data source is not subject to frequent updates.
Add to cache:
```js
ctx.cache.set((key: string), (value: string), (time: number)); // time: the cache expiry time in seconds
```
Access the cache:
```js
const value = await ctx.cache.get((key: string));
```
In this example: [/routes/zhihu/daily.js](https://github.com/DIYgod/RSSHub/blob/master/routes/zhihu/daily.js), the full text of each article is required resulting in many requests being sent. The update frequency for this source is known (daily), we can safely set the cache to a day to avoid wasting resources.
### Data
Assign the acquired data to ctx.state.data, the middleware [template.js](https://github.com/DIYgod/RSSHub/blob/master/middleware/template.js) will then process the data and render the RSS output [/views/rss.art](https://github.com/DIYgod/RSSHub/blob/master/views/rss.art), the list of parameters:
```js
ctx.state.data = {
title: '', // The feed title
link: '', // The feed link
description: '', // The feed description
language: '', // The language of the channel
item: [
// An article of the feed
{
title: '', // The article title
author: '', // Author of the article
category: '', // Article category
// category: [''], // Multiple category
description: '', // The article summury or content
pubDate: '', // The article publishing datetime
guid: '', // The article unique identifier, optional, default to the article link below
link: '', // The article link
},
],
};
```
#### Podcast feed
Used for audio type feed, these **additional** datas can make your podcast subscribeable:
```js
ctx.state.data = {
itunes_author: '', // The channel's author, you must fill this data.
itunes_category: '', // Channel category
image: '', // Channel's image
item: [
{
itunes_item_image: '', // The item image
enclosure_url: '', // The item's audio link
enclosure_length: '', // The audio length, the unit is seconds.
enclosure_type: '', // 'audio/mpeg' or 'audio/x-m4a' or others
},
],
};
```
#### BT feed
Used for download type feed, these **additional** datas can make your BT client subscribeable and can auto download:
```js
ctx.state.data = {
item: [
{
enclosure_url: '', // Magnet URI
enclosure_length: '', // The audio length, the unit is seconds, optional
enclosure_type: 'application/x-bittorrent', // Fixed to 'application/x-bittorrent'
},
],
};
```
</details>
---
## Join the discussion
1. [Telegram Group](https://t.me/rsshub)
2. [GitHub Issues](https://github.com/DIYgod/RSSHub/issues)
+25 -14
View File
@@ -36,7 +36,7 @@ sidebar: auto
});
const data = response.data.data; // response.data 为 HTTP GET 请求返回的数据对象
//这个对象中包含了数组名为 data,所以 response.data.data 则为需要的数据
// 这个对象中包含了数组名为 data,所以 response.data.data 则为需要的数据
```
返回的数据样例之一(response.data.data[0]):
@@ -57,7 +57,7 @@ sidebar: auto
}
```
对数据进行进一步处理,生成符合 RSS 规范的对象,把获取的标题、链接、描述、发布时间等数据赋值给 ctx.state.data,[生成 RSS](#生成-rss):
对数据进行进一步处理,生成符合 RSS 规范的对象,把获取的标题、链接、描述、发布时间等数据赋值给 ctx.state.data, [生成 RSS 源](#生成-rss-源):
```js
ctx.state.data = {
@@ -104,7 +104,7 @@ sidebar: auto
```js
const $ = cheerio.load(data); // 使用 cheerio 加载返回的 HTML
const list = $('.note-list li').get();
// 使用 cheerio 选择器,选择 class="note-list" 下的所有 <li> 元素,返回 cheerio node 对象数组
// 使用 cheerio 选择器,选择 class="note-list" 下的所有 "li"元素,返回 cheerio node 对象数组
// cheerio get() 方法将 cheerio node 对象数组转换为 node 对象数组
// 注:每一个 cheerio node 对应一个 HTML DOM
@@ -163,8 +163,8 @@ sidebar: auto
guid: itemUrl,
};
// 使用tryGet方法从缓存获取内容。
// 当缓存中无法获取到链接内容的时候,则使用load方法加载文章内容。
// 使用 tryGet() 方法从缓存获取内容
// 当缓存中无法获取到链接内容的时候,则使用 load() 方法加载文章内容
const other = await caches.tryGet(itemUrl, async () => await load(itemUrl), 3 * 60 * 60);
// 合并解析后的结果集作为该篇文章最终的输出结果
@@ -258,6 +258,8 @@ sidebar: auto
// 注:由于此路由只是起到一个新专栏上架提醒的作用,无法访问付费文章,因此没有文章正文
```
---
#### 使用缓存
所有路由都有一个缓存,全局缓存时间在 `lib/config.js` 里设定,但某些接口返回的内容更新频率较低,这时应该给这些数据设置一个更长的缓存时间。
@@ -280,7 +282,7 @@ const value = await ctx.cache.get((key: string));
```js
const key = 'daily' + story.id; // story.id 为知乎日报返回的文章唯一识别符
ctx.cache.set(key, item.description, 24 * 60 * 60); // 设置缓存时间为 24小时 * 60分钟 * 60秒 = 86400秒/1天
ctx.cache.set(key, item.description, 24 * 60 * 60); // 设置缓存时间为 24小时 * 60分钟 * 60秒 = 86400秒 = 1天
```
当同样的请求被发起时,优先使用未过期的缓存:
@@ -297,7 +299,9 @@ if (value) {
}
```
#### 生成 RSS
---
#### 生成 RSS 源
获取到的数据赋给 ctx.state.data, 然后数据会经过 [template.js](https://github.com/DIYgod/RSSHub/blob/master/lib/middleware/template.js) 中间件处理,最后传到 [/lib/views/rss.art](https://github.com/DIYgod/RSSHub/blob/master/lib/views/rss.art) 来生成最后的 RSS 结果,每个字段的含义如下:
@@ -343,7 +347,7 @@ ctx.state.data = {
};
```
##### BT 源
##### BT/磁力源
用于下载类 RSS,**额外**添加这些字段能使你的 RSS 被 BT 客户端识别并自动下载:
@@ -359,15 +363,17 @@ ctx.state.data = {
};
```
---
### 步骤 2: 添加脚本路由
在 [/lib/router.js](https://github.com/DIYgod/RSSHub/blob/master/lib/router.js) 里添加路由:
在 [/lib/router.js](https://github.com/DIYgod/RSSHub/blob/master/lib/router.js) 里添加路由
#### 举例
1. [bilibili/bangumi](https://github.com/DIYgod/RSSHub/blob/master/lib/routes/bilibili/bangumi.js)
| 类型 | 代码 |
| 名称 | 说明 |
| -------------------------- | ---------------------------------------------------------------------------------- |
| 路由 | `/bilibili/bangumi/:seasonid` |
| 数据来源 | bilibili |
@@ -378,9 +384,9 @@ ctx.state.data = {
| 脚本路径 | `./routes/bilibili/bangumi` |
| lib/router.js 中的完整代码 | `router.get('/bilibili/bangumi/:seasonid', require('./routes/bilibili/bangumi'));` |
1. [github/issue](https://github.com/DIYgod/RSSHub/blob/master/lib/routes/github/issue.js)
2. [github/issue](https://github.com/DIYgod/RSSHub/blob/master/lib/routes/github/issue.js)
| 类型 | 代码 |
| 名称 | 说明 |
| -------------------------- | ---------------------------------------------------------------------------- |
| 路由 | `/github/issue/:user/:repo` |
| 数据来源 | github |
@@ -391,9 +397,9 @@ ctx.state.data = {
| 脚本路径 | `./routes/github/issue` |
| lib/router.js 中的完整代码 | `router.get('/github/issue/:user/:repo', require('./routes/github/issue'));` |
1. [embassy](https://github.com/DIYgod/RSSHub/blob/master/lib/routes/embassy/index.js)
3. [embassy](https://github.com/DIYgod/RSSHub/blob/master/lib/routes/embassy/index.js)
| 类型 | 代码 |
| 名称 | 说明 |
| -------------------------- | ---------------------------------------------------------------------------- |
| 路由 | `/embassy/:country/:city?` |
| 数据来源 | embassy |
@@ -404,6 +410,8 @@ ctx.state.data = {
| 脚本路径 | `./routes/embassy/index` |
| lib/router.js 中的完整代码 | `router.get('/embassy/:country/:city?', require('./routes/embassy/index'));` |
---
### 步骤 3: 添加脚本文档
1. 更新 [文档 (/docs/README.md) ](https://github.com/DIYgod/RSSHub/blob/master/docs/README.md), 可以执行 `npm run docs:dev` 查看文档效果
@@ -478,6 +486,9 @@ ctx.state.data = {
1. 执行 `npm run format` 自动标准化代码格式,提交代码, 然后提交 pull request
---
## 参与讨论
1. [Telegram 群](https://t.me/rsshub)
2. [GitHub Issues](https://github.com/DIYgod/RSSHub/issues)