# Question regarding keyword search

**URL:** <https://sleuthkit.discourse.group/t/question-regarding-keyword-search/369>\
**Category:** Autopsy Help\
**Created:** [November 18, 2019, 12:17pm UTC](https://sleuthkit.discourse.group/t/question-regarding-keyword-search/369 "2019-11-18T12:17:49Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Tic-Tac](https://avatars.discourse-cdn.com/v4/letter/t/977dab/32.png) [@Tic-Tac](https://sleuthkit.discourse.group/u/Tic-Tac)\
**Post date:** [November 18, 2019, 12:17pm UTC](https://sleuthkit.discourse.group/t/question-regarding-keyword-search/369/1 "2019-11-18T12:17:49Z")

</div>

Dear forensics enthusiasts!

Soon I will have to go out in the field and perform live forensics. I will have to perform keyword search and look for various documents.

I have a live Linux system on a USB stick with various utilities, including the latest version of Autopsy and I’m wondering how exactly does the keyword search works.

Time will be of an essence and thus I’m wondering how to approach this task. I intend to run the “Indexing and keyword search” module but I am wondering will it provide me with keyword hits from .docx and similar documents if I do not run the “Embedded file extractor” module beforehand?

Looking forward to your insight and thoughts on how to approache this task in the best possible way 🙂

---

<div class="post-metadata">

**Author:** ![downey](https://avatars.discourse-cdn.com/v4/letter/d/a87d85/32.png) [@downey](https://sleuthkit.discourse.group/u/downey)\
**Post date:** [November 18, 2019, 9:18pm UTC](https://sleuthkit.discourse.group/t/question-regarding-keyword-search/369/2 "2019-11-18T21:18:17Z")

</div>

> …will it provide me with keyword hits from .docx and similar documents if I do not run the “Embedded file extractor” module beforehand?

Yes.

There are way too many unknowns to give you any useful feedback on overall approach. The only general feedback I can provide is probably obvious…test your planned methodology and configuration in a lab environment before doing it in the ‘field’.

---

<div class="post-metadata">

**Author:** ![apriestman](https://yyz2.discourse-cdn.com/free1/user_avatar/sleuthkit.discourse.group/apriestman/32/24_2.png) [@apriestman](https://sleuthkit.discourse.group/u/apriestman)\
**Post date:** [November 19, 2019, 1:36pm UTC](https://sleuthkit.discourse.group/t/question-regarding-keyword-search/369/3 "2019-11-19T13:36:34Z")

</div>

> [@Tic-Tac](#):
>
> I intend to run the “Indexing and keyword search” module but I am wondering will it provide me with keyword hits from .docx and similar documents if I do not run the “Embedded file extractor” module beforehand?

To expand a bit on the previous answer - the Embedded File Extractor module most commonly extracts files from archives (.zip, .rar, etc) and images from documents (.docx and others). All of these extracted files are then processed by any ingest modules you have selected. So you don’t need to worry about order - just run all the ingest modules you need. And if you choose not to run the Embedded File Extractor module you’ll still be able to do run keyword search on the documents.

If you haven’t, I would suggest reading the help page on Keyword Search and practicing with it before you go into the field if at all possible. You can add some files and folders on your machine as a logical files data source to test how the searching works.

[http://sleuthkit.org/autopsy/docs/user-docs/4.13.0/keyword\_search\_page.html](http://sleuthkit.org/autopsy/docs/user-docs/4.13.0/keyword_search_page.html)  
[http://sleuthkit.org/autopsy/docs/user-docs/4.13.0/ds\_page.html#ds\_log](http://sleuthkit.org/autopsy/docs/user-docs/4.13.0/ds_page.html#ds_log)

---

<div class="post-metadata">

**Author:** ![downey](https://avatars.discourse-cdn.com/v4/letter/d/a87d85/32.png) [@downey](https://sleuthkit.discourse.group/u/downey)\
**Post date:** [November 19, 2019, 9:23pm UTC](https://sleuthkit.discourse.group/t/question-regarding-keyword-search/369/4 "2019-11-19T21:23:45Z")

</div>

You might also find the following documentation useful for the scenario you describe.

[http://sleuthkit.org/autopsy/docs/user-docs/4.13.0//triage\_page.html](http://sleuthkit.org/autopsy/docs/user-docs/4.13.0//triage_page.html)

---

<div class="post-metadata">

**Author:** ![Tic-Tac](https://avatars.discourse-cdn.com/v4/letter/t/977dab/32.png) [@Tic-Tac](https://sleuthkit.discourse.group/u/Tic-Tac)\
**Post date:** [December 10, 2019, 8:24am UTC](https://sleuthkit.discourse.group/t/question-regarding-keyword-search/369/5 "2019-12-10T08:24:38Z")

</div>

Thank you everyone for the valuable insights and recommendations, you’ve been very helpful, will definitely do some small experiments and tests 🙂
