ascii character - removing chars from string

bruce · Jul 4, 2006

hi...

i'm running into a problem where i'm seeing non-ascii chars in the parsing
i'm doing. in looking through various docs, i can't find functions to
remove/restrict strings to valid ascii chars.

i'm assuming python has something like

valid_str = strip(invalid_str)

where 'strip' removes/strips out the invalid chars...

any ideas/thoughts/pointers...

thanks

-bruce

bearophileHUGS · Jul 4, 2006

bruce:

valid_str = strip(invalid_str)
where 'strip' removes/strips out the invalid chars...

This isn't short but it is fast:
import string
valid_chars = string.lowercase + string.uppercase + \
string.digits +
"""|!'\\"£$%&/()=?^*é§_:;>+,.-<\n \t"""
all_chars = "".join(map( chr, range(256)) )
comp_valid_chars = "".join( set(all_chars).difference(valid_chars) )
print "test string".translate(all_chars, comp_valid_chars)

Shorter and a bit slower alternative:
import string
valid_chars_set = set(string.lowercase + string.uppercase
+ string.digits +
"""|!'\\"£$%&/()=?^*é§_:;>+,.-<\n \t""")
print filter(lambda c: c in valid_chars_set, "test string")

You can add the chars you want to the string of accepted ones.

Bye,
bearophile

John Machin · Jul 4, 2006

hi...

i'm running into a problem where i'm seeing non-ascii chars in the parsing
i'm doing. in looking through various docs, i can't find functions to
remove/restrict strings to valid ascii chars.

It's possible that you would be better off handling those characters in
some fashion other than blowing them away. What are the characters that
you are seeing, and what is the problem that they are causing you?

Rune Strand · Jul 4, 2006

bruce said:
hi...

i'm running into a problem where i'm seeing non-ascii chars in the parsing
i'm doing. in looking through various docs, i can't find functions to
remove/restrict strings to valid ascii chars.

i'm assuming python has something like

valid_str = strip(invalid_str)

where 'strip' removes/strips out the invalid chars...

any ideas/thoughts/pointers...

If you're able to define the invalid_chars, the most convenient is
probably to use the strip() method:'def'

Simon Forman · Jul 4, 2006

bruce said:
hi...

i'm running into a problem where i'm seeing non-ascii chars in the parsing
i'm doing. in looking through various docs, i can't find functions to
remove/restrict strings to valid ascii chars.

i'm assuming python has something like

valid_str = strip(invalid_str)

where 'strip' removes/strips out the invalid chars...

any ideas/thoughts/pointers...

thanks

-bruce

You might be able to use the translate() and maketrans() string
methods. See
http://groups.google.ca/group/comp...._frm/thread/261516dd06ee32e6/9a02a21c95bd0ec9
second-to-last post for an example.

Simon Forman · Jul 4, 2006

bruce said:
hi...

update. i'm getting back html, and i'm getting strings like " foo  "
which is valid HTML as the ' ' is a space.

&, n, b, s, p, ; Those are all ascii characters.

i need a way of stripping/removing the ' ' from the string

the   needs to be treated as a single char...

text = "foo cat  "

ie ok_text = strip(text)

ok_text = "foo cat"

Do you really want to remove those html entities? Or would you rather
convert them back into the actual text they represent? Do you just
want to deal with  's? Or maybe the other possible entities that
might appear also?

Check out htmlentitydefs.entitydefs (see
http://docs.python.org/lib/module-htmlentitydefs.html) it's kind of
ugly looking so maybe use pprint to print it:
{'AElig': 'Æ',
'Aacute': 'Á',
'Acirc': 'Â',
..
..
..
'nbsp': '\xa0',
..
..
..
etc...

HTH,
~Simon

"You keep using that word. I do not think it means what you think it
means."
-Inigo Montoya, "The Princess Bride"

Simon Forman · Jul 4, 2006

bruce said:
simon...

the ' ' is not to be seen/viewed as text/ascii.. it's a representation
of a hex 'u\xa0' if i recall...

Did you not see this part of the post that you're replying to?

'nbsp': '\xa0',

My point was not that '\xa0' is an ascii character... It was that your
initial request was very misleading:

"i'm running into a problem where i'm seeing non-ascii chars in the
parsing i'm doing. in looking through various docs, i can't find
functions to remove/restrict strings to valid ascii chars."

That's why you got three different answers to the wrong question.

You weren't "seeing non-ascii chars" at all. You were seeing ascii
representations of html entities that, in the case of ' ', happen
to represent non-ascii values.

i'm looking to remove or replace the insances with a ' ' (space)

Simplicity:

s.replace(' ', ' ')

~Simon

"You keep using that word. I do not think it means what you think it
means."
-Inigo Montoya, "The Princess Bride"

-bruce

-----Original Message-----
From: [email protected]
[mailto[email protected]]On Behalf
Of Simon Forman
Sent: Monday, July 03, 2006 7:17 PM
To: (e-mail address removed)
Subject: Re: ascii character - removing chars from string

hi...

update. i'm getting back html, and i'm getting strings like " foo  "
which is valid HTML as the ' ' is a space.

Click to expand...

&, n, b, s, p, ; Those are all ascii characters.

i need a way of stripping/removing the ' ' from the string

the   needs to be treated as a single char...

text = "foo cat  "

ie ok_text = strip(text)

ok_text = "foo cat"

Click to expand...

Do you really want to remove those html entities? Or would you rather
convert them back into the actual text they represent? Do you just
want to deal with  's? Or maybe the other possible entities that
might appear also?

Check out htmlentitydefs.entitydefs (see
http://docs.python.org/lib/module-htmlentitydefs.html) it's kind of
ugly looking so maybe use pprint to print it:
{'AElig': 'Æ',
'Aacute': 'Á',
'Acirc': 'Â',
.
.
.
'nbsp': '\xa0',
.
.
.
etc...

HTH,
~Simon

"You keep using that word. I do not think it means what you think it
means."
-Inigo Montoya, "The Princess Bride"

Ascii to Unicode.	4	Jul 28, 2010
strip() using strings instead of chars	4	Jul 11, 2008
PEP 3131: Supporting Non-ASCII Identifiers	399	May 13, 2007
find and remove "\" character from string	2	Sep 15, 2007
Lowest addressed character in array	8	Dec 14, 2010
DBD::Oracle, Unicode, non-UTF8-non-ASCII strings	0	Jul 23, 2009
Elementary string-parsing	16	Feb 4, 2008
string/list comparison	0	Jul 6, 2006

ascii character - removing chars from string

bruce

bearophileHUGS

John Machin

Rune Strand

Simon Forman

Simon Forman

Simon Forman

Ask a Question

Similar Threads

Members online

Forum statistics

Latest Threads