fix(core): invalid feed fields (#9286)

Signed-off-by: Rongrong <15956627+Rongronggg9@users.noreply.github.com>
This commit is contained in:
Rongrong
2022-03-22 02:13:15 +08:00
committed by GitHub
parent fe775f5198
commit c48ca6bd5b
4 changed files with 40 additions and 1 deletions
+10
View File
@@ -322,6 +322,16 @@ ctx.state.data = {
};
```
::: warning Warning
`title`, `subtitle` (only for atom), `author` (only for atom), `item.title`, and `item.author` should not contain linebreaks, consecutive white spaces, or start/end with white space(s).
Most RSS readers will automatically trim them, so they make no sense. However, some readers may not process them properly, so we will trim them before outputting to ensure these fields contain no linebreaks, consecutive white spaces, or start/end with white space(s).
If the route you are writing can not tolerate these trimmings, you should consider change the format of these fields.
In addition, although other fields will not be forced trimmed, you should also try to avoid violations of the above rules. Especially when using Cheerio to extract web pages, you need to keep in mind that Cheerio will retain wraps and indentation. In particular, for `item.description`, any intended linebreaks should be converted to `<br>`, otherwise the RSS reader is likely to trim them; especially if you extract the RSS feed from JSON, the JSON returned by the source website is very likely to contain linebreaks that need to be displayed, so it must be converted in this case.
:::
##### Podcast feed
Used for audio feed, these **additional** data are in accordance with many podcast players' subscription format:
+10
View File
@@ -323,6 +323,16 @@ ctx.state.data = {
};
```
::: warning 注意
`title`, `subtitle` (仅适用于 atom 输出), `author` (仅适用于 atom 输出), `item.title`, `item.author` 不应该包含换行、多于一个的连续空字符,或以空字符开头 / 结尾。\
多数 RSS 阅读器会自动为上述字段修剪空字符,所以这些空字符没有意义。但是,某些阅读器也许不能正确处理它们,因此,我们会在最终输出前修剪上述字段,确保不含有换行或多于一个空字符,也不以空字符开头或结尾。\
如果你要编写的路由在上述字段不能容忍空字符修剪,你应该考虑变换一下这些字段的格式。
另外,虽然其它字段不会经过强制空字符修剪,但你也应该尽量避免违反上述规则。尤其是使用 cheerio 提取网页元素或文本时,需要时刻谨记 cheerio 会保留换行和缩进。特别地,对于 `item.description` ,任何预期之内的换行都应被转换为 `<br>` ,否则 RSS 阅读器很可能将它修剪;尤其如果你从 JSON 提取 RSS 源,目标网站返回的 JSON 很有可能含有需要显示的换行,这时候就一定要进行转换。
:::
##### 播客源
用于音频类 RSS,**额外**添加这些字段能使你的 RSS 被泛用型播客软件订阅:
+8 -1
View File
@@ -2,6 +2,7 @@ const art = require('art-template');
const path = require('path');
const config = require('@/config').value;
const typeRegex = /\.(atom|rss|debug\.json)$/;
const { collapseWhitespace } = require('@/utils/common-utils');
module.exports = async (ctx, next) => {
if (ctx.headers['user-agent'] && ctx.headers['user-agent'].includes('Reeder')) {
@@ -40,10 +41,14 @@ module.exports = async (ctx, next) => {
}
if (ctx.state.data) {
ctx.state.data.title = collapseWhitespace(ctx.state.data.title);
ctx.state.data.subtitle = collapseWhitespace(ctx.state.data.subtitle);
ctx.state.data.author = collapseWhitespace(ctx.state.data.author);
ctx.state.data.item &&
ctx.state.data.item.forEach((item) => {
if (item.title) {
item.title = item.title.trim();
item.title = collapseWhitespace(item.title);
// trim title length
for (let length = 0, i = 0; i < item.title.length; i++) {
length += Buffer.from(item.title[i]).length !== 1 ? 2 : 1;
@@ -54,6 +59,8 @@ module.exports = async (ctx, next) => {
}
}
item.author = collapseWhitespace(item.author);
if (item.itunes_duration && ((typeof item.itunes_duration === 'string' && item.itunes_duration.indexOf(':') === -1) || (typeof item.itunes_duration === 'number' && !isNaN(item.itunes_duration)))) {
item.itunes_duration = +item.itunes_duration;
item.itunes_duration =
+12
View File
@@ -6,6 +6,18 @@ const toTitleCase = (str) =>
.map((word) => word.replace(word[0], word[0].toUpperCase()))
.join(' ');
const rWhiteSpace = /\s+/;
const rAllWhiteSpace = /\s+/g;
// collapse all whitespaces into a single space (like "white-space: normal;" would do), and trim
const collapseWhitespace = (str) => {
if (str && rWhiteSpace.test(str)) {
return str.replace(rAllWhiteSpace, ' ').trim();
}
return str;
};
module.exports = {
toTitleCase,
collapseWhitespace,
};