0
187 次浏览

[reposting from another list - TCOP]

I occasionally concatenate USFM files into a complete Bible text file in order to process them outside of Paratext. This consists of all of the files copied into one text file in sequential order.

One of the cumbersome manual steps in this process has been to separate them back into individual book files for long term storage.

Today, when I am importing a processed file back into Paratext, I haven’t separated the file properly: Each book file exists, but the text file Mark also contains John to Revelation instead of ending with Mark 16. Paratext recognized this as “multiple files contain information about the same book.”

When I try to import only the concatenated file which contains the books Mark-Rev, I get all these books in Paratext, with no apparent loss or added information.

This is foundationally different than the USFM I know, where 1 Bible book == 1 text file. Is this {Bible Book Aggregation|File concatenation} a feature? Does Paratext support importing concatenated files officially?

Taking it one step further: my post processing usually includes file concatenation. Aggregating is to a degree part of our local implementation of USFM. Is book aggregation actually part of the USFM specification? (and if not, could it be?)

I might owe someone a coffee if this is true … 1 bible = 1 text file would save so so much file manipulation time for me.

Thanks,

Michael Hart
Senior Publishing Services Specialist
Bible League International

较早的帖子 - 以原始语言显示
Paratext (149 分) 发布 | 187 次浏览

2 个回答

+1
最佳答案

I believe that an attempt was made to many years ago to allow importing multiple USFM books from a single file. Most people don’t know about this and it is (I believe) it is rarely done so I can’t guarantee that you won’t find some way to break it.

Another thing that is probably mostly unknown is that you can import USX.

较早的帖子 - 以原始语言显示
发布
0

After I posted this I did more research and discovered the linux/unix command csplit provides most of the functionality I need to restore a concatenated USFM file into it’s component pieces. That is, the linux command

$csplit -k /\\id\ / {1,100} *.sfm

breaks a joined file into SFM compatible files that still need naming properly.

Is there a windows equivalent to csplit? That is a known way to split a file by its contents from the command line?

Note that I"m working with American English files. In the past, I’ve had problem with Chinese characters being corrupted during the join step using the linux command line in some systems, even those that claimed POSIX compliance. If you try this at home and are working with unicode < Ux2100, I’d like to hear how it works.

较早的帖子 - 以原始语言显示
(149 分) 发布

相关问题

0
3 个回答 286 次浏览
I see that Paratext 8.0.100.1 comes with USFM v. 2.502. According to http://ubsicap.github.io/usfm/ it looks as if 3. ... version. Why is that version not used for PT 8.0.100.1?
anon716631 346 发布 提问于 四月 11, 2017
0
1 个回答 233 次浏览
各位同事,你们好 我正在帮助一个团队处理他们的 \r 平行参考标记,这些标记通常打印在小节标题下方 在某些情况下,他们希望这些标记包含更多信息,在小节顶部显示该平行参考具体适用于哪些经文(即目标交叉引用的来源) 例如: \s1 ... 改 这是最好的方法吗,还是有更好的标准做法?如果这些信息都放在页面底部,我很乐意使用标准的交叉引用标记 谢谢 Phil
anon913937 124 发布 提问于 十一月 20, 2023
0
2 个回答 45 次浏览
我需要将一个 USX 文件转换为 USFM。我看到 Paratext 有反向操作的选项,但没看到从 USX 到 USFM 的选项。我需要这样做,因为 Proskomma(用于 Scripture App Builder 构建 PWA)拒绝了圣经中的 12 卷书,原因是它们不符合标准格式。我希望将它们转换为 USFM 后,Proskomma 处理起来会更轻松。
yeti 109 发布 提问于 三月 19
0
1 个回答 34 次浏览
我们在脚注交叉引用中使用了 /xt,设置为使用书籍的“短名”。但在词汇表中,我想使用书籍名称的缩写,是否有不同的 USFM 代码可以实现这一点?
Lenice 180 发布 提问于 十月 7, 2025
0
0 个回答 51 次浏览
目前是否已经存在针对 USFM 的语法高亮工具? 这些工具可能以以下形式存在: 模式匹配器 解析器 语义高亮 我想知道在其他 USFM 编辑器中提供此功能有多困难。
Mercado 203 发布 提问于 十月 24, 2024
Welcome to Support Bible, where you can ask questions and receive answers from other members of the community.
And let us consider how we may spur one another on toward love and good deeds, not giving up meeting together, as some are in the habit of doing, but encouraging one another—and all the more as you see the Day approaching.
Hebrews 10:24-25
3,045 个问题
6,005 个回答
5,671 条评论
2,026 位用户