best technique for detecting charset/character encoding of RSS feeds

D

Daniel Choi

I have a Ruby application that fetches RSS and Atom feeds. I've tried
using the chardet gem (UniversalDetector) to figure out what the
character encoding of each feed is. But for some strange this library
thinks a lot of feeds are EUC-KR (Korean) when they plainly aren't.

Can anyone suggest a better way to find the encoding of RSS and Atom
feeds (e.g. via BOM detection, etc.) with Ruby?
 

Ask a Question

Want to reply to this thread or ask your own question?

You'll need to choose a username for the site, which only take a couple of moments. After that, you can post your question and our members will help you out.

Ask a Question

Members online

Forum statistics

Threads
473,769
Messages
2,569,582
Members
45,071
Latest member
MetabolicSolutionsKeto

Latest Threads

Top