"WordList - ExampleSet conversion problem when using Bayes Classifier - text cat."

Question

Hey Guys, I am working on a small project, and basically want to categorize text using a Naive Bayes classifier. What I have done - created a process that trains the bayes classifier (using processed data) and serializes the classifier using write Model. Furthermore, I have also stored the word list that was generated by the process documents from data (by using Wordlist to Data and storing it as an ARFF). The process I am working on, and which I'm having problems with is the model applier to the data. I have a file which has a single line of text (the document to be categorized). I feed it into process documents, and have to also feed the aforementioned word list. So I load the word list by read ArFF and connect it to the wordlist input port on the process documents. However an error is thrown saying 'Expected WordList but received ExampleSet'. How do i generate a WordList from the data? Or is the way I am storing it incorrect. Thanks for any help! I really appreciate it! The following is my XML:

MariusHelf · Answer

Hi m1cros,

in general, your process setup looks fine. However, you should strongly consider to store all objects you create in the repository using the Store operator or the process context. That way you can retrieve them later with the Retrieve operator for use in other processes. You can store basically anything in the repository: ExampleSets, Models, Performances, Wordlists etc. I don't know how you stored your wordlist in a file, but storing it in the repository also guarantees that the type of the object is remembered correctly.

You may also consider to store your data there, because repository access is usually faster that Read CSV, plus you can store metadata like attribute roles etc. in the repository and don't have to preprocess your data on each access.

Cheers, Marius