SoraTranslator is a tool designed to extract, translate, and reinsert text in Galgames (Japanese visual novels). This tool primarily automates the translation process for these games. The GUI provides a easy-to-use interface for correcting manual or compare the translation results.
SoraTranslator operates by integrating with specific Galgames, focusing on the extraction, translation, and reintegration of game text. It's structured to handle various aspects of this process, from accessing original game resources to generating translated files ready for gameplay.
It focuses on the improvement of the translation quality by human intervention. The translation is done by the AI translator selected, but the user can choose to modify the translation if it is not satisfactory by modifying the text in an Excel-like table powered by Handsontable.
The build-in translator focuses on saving tokens, by providing context by a specific separation method (now only works with the 16k model). Any misalignment of the text will be recorded and the translator can easily fix it. Other translators are also supported, like the GalTransl translator, which is included in the program.
The translator module can also be used for other purposes, but some functions (like keeping the same amount of blocks in the translation) might not be needed in other cases.
Figure: The start page of the program.
Figure: Translation page of the program
The project is organized as follows:
SoraTranslator/
|─── backend/
| |── Integrators/
│ ├── GameName/
│ ├── Extractor/
│ ├── Parser/
| |── Translators/
|─── frontend/
| |...
GameFolder
|─── SoraTranslator/
│ ├── OriginalFiles/
│ ├── RawText/
│ ├── Text/
│ ├── TranslatedFiles/
Each directory is designed to handle a different stage of the translation process:
- Integrators: Contains tools for extracting and parsing game text.
- Translators: Dedicated to the translation process.
- GameFolder/SoraTranslator: Stores the original game files, extracted text, translated text, and the final translated game files. This folder is chosen by the user.
- Python >= 3.8
- Node.js >= 20.0.0
- API key for your selected endpoint (OpenAI/Gemini/Grok/etc.)
If you use the released version, the requirements may not be necessary.
- Clone the repository, or download the latest release zip and unzip it.
- If using source mode, build frontend executable (
npm run electron:buildinsidefrontend) and package withpackage.bat. - For portable releases, use the deterministic layout documented in docs/release-layout.md.
- Unzip the release folder.
- Run
setup-config.batonce and enter endpoint/model/API key. - Run
run.bat.
- Install backend/frontend dependencies as needed.
- Run
run.bat --dev.
The launcher now:
- auto-detects Python (
py -3fallback topython) - creates
backend/.venvif missing - skips
pip installwhen requirements and Python version are unchanged - waits for backend health before launching frontend
- cleans backend process on frontend exit
Python 3 not found: install Python 3.8+ and ensurepyorpythonis inPATH.Backend did not become ready: checkbackend/backend.logand verify local firewall/antivirus rules.- Endpoint auth errors: rerun
setup-config.batand verify the key for the selected endpoint. - To run launcher checks without frontend, use
run.bat --smoke. - For quick local checks without package reinstall work, use
run.bat --smoke --skip-install. - Launcher execution details are written to
log.txtin the app root (mirrored tolauncher.log.txtfor compatibility). - Dev/release runs are unlimited by default; set
--max-runtime-seconds <N>if you want an enforced hard stop. - Smoke runs keep a short default runtime guard for launcher checks.
- Release mode uses backend port
5000by default for compatibility; override with--backend-port <N>. - Startup/install wait time is independently bounded; tune with
--startup-timeout-seconds <N>. - Packaging logs are written to
package.log.txt.
- Do not commit plaintext API keys.
- Store runtime keys in
.env(created by setup) and reference them withENV:<KEY_NAME>in config files. config.template.json,translators.json, and docs are secret-safe defaults only.- Legacy plaintext keys are still read for compatibility, but migration to env-backed values is recommended.
A usage tutorial video can be found at here.
To utilize SoraTranslator for different games, particularly from various companies, custom integrators must be written to handle the specific file formats and structures of each game.
The extraction could be empty if you already have the text extracted from the game.
The parser, no matter how it works, should output the text in a uniform format (in .csv files), with the following structure:
You don't have to use the data structures written in this program. But your game definition file must contain the functions declared in backend/game.py.
The comma is better for illustration, use "\t" in the real file.
block1_name,original_speaker,original_text,translated_speaker,translated_text,is_translated,translation_date,translation_method
block2_name,original_speaker,original_text,translated_speaker,translated_text,is_translated,translation_date,translation_method
...
The first 5 parts are necessary, so every line should have at least 5 parts (4 "\t"s) for the translation engine to read the text.
The program will soon add support for .json format, which will be more flexible.
The translation process is handled by the Translator module, which is designed to function separately from the integrator. This allows the extension of the usage of the program. For example, you can put the text of Doujinshi in the Texts folder and translate it without the need to integrate it into a game. However, the above format is still required.
The translator will translate the text file in '.csv' and replace the text in the 4th and 5th columns as shown above. The translated text will be saved in the same file. The translation method and date will also be recorded.
There can be multiple translators for the program. However, the program is designed to work with GPT API or manual translation.
Usage of the GalTransl translator is suggested, it is included now in the program, thanks to GalTransl!
- Extraction: The integrator accesses the
OriginalFilesdirectory (the best way is to put a "file_path.py" inside the folder recording all files and information about the game), extracting and parsing game text into a readable format (.csv) in theTextsfolder. (There will also be aRawTextfolder for the raw text extracted from the game, this allows easier "roll back" if you did something wrong.) - Translation: Utilizing the Translator module (at the front end), the text in
Textsis then translated, with the results stored in the same file. - Reintegration: Finally, the integrator takes the translated text in
Textsand repackages it into the game's raw text format, saving these files in theTranslatedFilesdirectory for use in the game.
- Translate the speakers and options
- Add support for more games
- Add support for MTools
- Add support for translating doujinshi
- OpenAI for the possibility, which helps me to translate what I've been wanting to play for centuries
- Handsontable for the table
- ONScripter-EN-Steam for the tools around NScripter
- XP3Unpacker for the tools around Kirikiri
- GalTransl