this post was submitted on 13 Nov 2023
12 points (80.0% liked)

Selfhosted

39980 readers
725 users here now

A place to share alternatives to popular online services that can be self-hosted without giving up privacy or locking you into a service you don't control.

Rules:

  1. Be civil: we're here to support and learn from one another. Insults won't be tolerated. Flame wars are frowned upon.

  2. No spam posting.

  3. Posts have to be centered around self-hosting. There are other communities for discussing hardware or home computing. If it's not obvious why your post topic revolves around selfhosting, please include details to make it clear.

  4. Don't duplicate the full text of your blog or github here. Just post the link for folks to click.

  5. Submission headline should match the article title (don’t cherry-pick information from the title to fit your agenda).

  6. No trolling.

Resources:

Any issues on the community? Report it using the report flag.

Questions? DM the mods!

founded 1 year ago
MODERATORS
 

I’m looking for a duplicate/similarity checker against a custom set of documents. This is possibly like a plagiarism checker, but with a custom reference (instead of everything that exists).

But I could not find a solution that can be selfhosted, and have some simple UI and capabilities like Turnitin. Any suggestions?

Thanks’

top 3 comments
sorted by: hot top controversial new old
[–] lemmy@linkopath.com 2 points 1 year ago

Maybe look into self hosted llm. I've used two recently to help analyze a large volume of books, by ingesting them into the data set, then chat with the bot for specifics. It worked pretty well but there are some limitations, such as token length and general hallucinations. But they both use citations of the data they used, so it helps to check their work.

PrivateGPT - https://github.com/imartinez/privateGPT

Llamaindex - https://github.com/run-llama/llama_index

Both have simple selfhosted webui or actual applications. So, in theory you should be able to ingest data and then see if it then matches any submissions you submit later. But I have not really tried it for this though, so it might not work.

[–] mojo@lemm.ee 1 points 1 year ago

This is very specific and niche without a good reason to exist, so I doubt this exists.

[–] AnonymousLemming@feddit.de 0 points 1 year ago

Grep. If you have the sources digitally and local, you can just run a search for a string.