开发者

Storing data to SequenceFile from Apache Pig

开发者 https://www.devze.com 2022-12-23 01:45 出处:网络
Apache Pig can load data from Hadoop seq开发者_运维问答uence files using the PiggyBank SequenceFileLoader:

Apache Pig can load data from Hadoop seq开发者_运维问答uence files using the PiggyBank SequenceFileLoader:

REGISTER /home/hadoop/pig/contrib/piggybank/java/piggybank.jar;

DEFINE SequenceFileLoader org.apache.pig.piggybank.storage.SequenceFileLoader();

log = LOAD '/data/logs' USING SequenceFileLoader AS (...)

Is there also a library out there that would allow writing to Hadoop sequence files from Pig?


It's just a matter of implementing a StoreFunc to do so.

This is possible now, although it will become a fair bit easier once Pig 0.7 comes out, as it includes a complete redesign of the Load/Store interfaces.

The "Hadoop expansion pack" Twitter is about to open source open-sourced at github, includes code for generating Load and Store funcs based on Google Protocol Buffers (building on Input/Output formats for same -- you already have those for sequence files, obviously). Check it out if you need examples of how to do some of the less trivial stuff. It should be fairly straightforward though.


This seemed to work for me. https://github.com/kevinweil/elephant-bird/pull/73

0

精彩评论

暂无评论...
验证码 换一张
取 消