Text Manipulation

tommy1234
tommy1234 New Altair Community Member
edited November 5 in Community Q&A
Hi there!
Its kinda hard to explain in words, so here's the example:
I would like to change all words such as: goood, good, goodddddd, gggooddd
to just good
Is there an operator that is suitable for this case in rapidminer?

Thanks in advanced!


Answers

  • MartinLiebig
    MartinLiebig
    Altair Employee
    regex and the replace operator is your friend, see attached example:

    Best,
    Martin

    <?xml version="1.0" encoding="UTF-8"?><process version="9.5.001"><br>  <context><br>    <input/><br>    <output/><br>    <macros/><br>  </context><br>  <operator activated="true" class="process" compatibility="9.5.001" expanded="true" name="Process"><br>    <parameter key="logverbosity" value="init"/><br>    <parameter key="random_seed" value="2001"/><br>    <parameter key="send_mail" value="never"/><br>    <parameter key="notification_email" value=""/><br>    <parameter key="process_duration_for_mail" value="30"/><br>    <parameter key="encoding" value="SYSTEM"/><br>    <process expanded="true"><br>      <operator activated="true" class="utility:create_exampleset" compatibility="9.5.001" expanded="true" height="68" name="Create ExampleSet" width="90" x="246" y="85"><br>        <parameter key="generator_type" value="comma separated text"/><br>        <parameter key="number_of_examples" value="100"/><br>        <parameter key="use_stepsize" value="false"/><br>        <list key="function_descriptions"/><br>        <parameter key="add_id_attribute" value="false"/><br>        <list key="numeric_series_configuration"/><br>        <list key="date_series_configuration"/><br>        <list key="date_series_configuration (interval)"/><br>        <parameter key="date_format" value="yyyy-MM-dd HH:mm:ss"/><br>        <parameter key="time_zone" value="SYSTEM"/><br>        <parameter key="input_csv_text" value="data&#10;goodddd&#10;goooodddd&#10;good"/><br>        <parameter key="column_separator" value=","/><br>        <parameter key="parse_all_as_nominal" value="false"/><br>        <parameter key="decimal_point_character" value="."/><br>        <parameter key="trim_attribute_names" value="true"/><br>      </operator><br>      <operator activated="true" class="replace" compatibility="9.5.001" expanded="true" height="82" name="Replace" width="90" x="514" y="85"><br>        <parameter key="attribute_filter_type" value="single"/><br>        <parameter key="attribute" value="data"/><br>        <parameter key="attributes" value=""/><br>        <parameter key="use_except_expression" value="false"/><br>        <parameter key="value_type" value="nominal"/><br>        <parameter key="use_value_type_exception" value="false"/><br>        <parameter key="except_value_type" value="file_path"/><br>        <parameter key="block_type" value="single_value"/><br>        <parameter key="use_block_type_exception" value="false"/><br>        <parameter key="except_block_type" value="single_value"/><br>        <parameter key="invert_selection" value="false"/><br>        <parameter key="include_special_attributes" value="false"/><br>        <parameter key="replace_what" value="g(o{2,})(d+)"/><br>        <parameter key="replace_by" value="good"/><br>      </operator><br>      <connect from_op="Create ExampleSet" from_port="output" to_op="Replace" to_port="example set input"/><br>      <connect from_op="Replace" from_port="example set output" to_port="result 1"/><br>      <portSpacing port="source_input 1" spacing="0"/><br>      <portSpacing port="sink_result 1" spacing="0"/><br>      <portSpacing port="sink_result 2" spacing="0"/><br>    </process><br>  </operator><br></process><br><br>




  • tommy1234
    tommy1234 New Altair Community Member
    Thank you for your answer @mschmitz
    I'm sorry but the attached example could not be seen.