Parse a TSV file_问答_开发者_运维开发者技术经验分享

开发者 https://www.devze.com 2022-12-22 15:05 出处：网络

I need to parse a file in TSV format (tab separated values). I use a regex to break down the file into each line, but I cannot find a satisfying one to parse each line.

相关专题：parsing regex

I need to parse a file in TSV format (tab separated values). I use a regex to break down the file into each line, but I cannot find a satisfying one to parse each line. For now I've come up this:

(?<g>("[^"]+")+|[^\t]+)

But it does not work if an item in the line has more than 2 consecutive double quotes.

Here's how the file is formatted: each element is separated by a tabulation. If an item contains a tab, it is encased with double quotes. If an item contains a double quote, it is doubled. But sometimes an element contains 4 conscutive double quotes, and the above regex splits the element into 2 different ones.

Examples:

item1ok "item""2""oK"

is correctly parsed into 2 elements: item1ok and item"2"ok (after trimming of the unnecessary quotes), but:

item1oK "item""""2oK"

is parsed into 3 elements: item1ok, item and "2ok (after trimming again).

Has anyone an idea how to make the regex fit this case? Or is there another solution to parse TSV simply? (I'm doing开发者_如何学Go this in C#).

You could use the TextFieldParser. This is technically a VB assembly, but you can use it even in C# by referencing the Microsoft.VisualBasic.FileIO assembly.

The example at the link above even shows using it on a tab-separated file.

Instead of trying to build your own CSV/TSV file parser (or using String.Split), I'd recommend you have a look at "Fast CSV Reader" or "FileHelpers library".

I'm using the first one, and am very happy with it (it supports any separator characters, e.g. comma, semicolon, tab).

Instead of using RegEx, maybe you could try the String.Split Method (Char[]) method.

I don't know C# but this should do the trick (in python)

txt = 'item1ok\t"item""2""oK"\titem1oK\t"item""""2oK"\tsomething else'
regex = '''
(?:                    # definition of a field
 "((?:[^"]|"")*)"   # either a double quoted field (allowing consecutive "")
 |                  # or
 ([^"]*)            # any character except a double quote
)                      # end of field
(?:$|\t)               # each field followed by a tab (except the last one)
'''
r = re.compile(regex, re.X)
# now find each match, and replace "" by " and remove trailing \t
# remove also the latest entry in the list (empty string)
columns = [t[0].replace('""', '"') if t[0] != '' else t[1].strip() for t in r.findall(txt)][:-1]
print columns
# prints: ['item1ok', 'item"2"oK', 'item1oK', 'item""2oK', 'something else']

Parse a TSV file

精彩评论

关注公众号

热门标签

图文推荐

Parse a TSV file

更多 问答 相关资讯：

精彩评论

关注公众号

热门标签

图文推荐

更多问答相关资讯：