Building Kitsune, a JavaScript obfuscator that respects its targets
Regex obfuscators break real code. I built Kitsune on the Babel AST instead — here's why, how the transforms work, and the gotchas around Google Apps Script that shaped its design.
- #javascript
- #obfuscation
- #babel
- #security
I keep a folder of obfuscation tools, and every JS one I’d tried had the same failure mode: regex. They hunt for string literals and identifiers with pattern matching, and the moment your code contains a template literal with a nested quote, or a string that happens to contain function, they mangle it into something that doesn’t run. For a throwaway script that’s annoying. For a Google Apps Script deployment that bills real users, it’s a liability.
So I built Kitsune (狐, fox) on the Babel toolchain instead: @babel/parser to read, @babel/traverse to walk, @babel/types to build nodes, @babel/generator to emit. Source-to-source, AST to AST. The code either transforms correctly or it doesn’t transform at all.
Why Google Apps Script changed the design
GAS is a weird deployment target. Your script is a single .gs file that Google’s runtime invokes through reserved entry points: doGet, onEdit, onOpen, and friends. It talks to globals like SpreadsheetApp and UrlFetchApp that exist nowhere in your file.
Every one of those names is a hard constraint on the obfuscator:
- Rename
doGetand the trigger never fires. The script deploys fine and does nothing, forever. - Rename
SpreadsheetAppand every call throws. The reference is resolved by the runtime, not the parser. - Encrypt the string
"SpreadsheetApp"and it’s even worse, because service lookups sometimes happen through bracket notation with string keys. The obfuscator can’t tell a service name from a label.
This is why Kitsune has a GOOGLE_RESERVED set that rename passes check before touching anything, plus --keep-funcs to preserve function and class names, and --keep-strings to leave string literals alone. They’re not conveniences, they’re the difference between an obfuscated script and a broken one.
The same set includes the obfuscator’s own runtime identifiers (__x, __k, __sd, __dl). The injected helpers use those names, so if your input code uses them, you get a collision. They’re reserved for the same reason.
The transforms
The default pipeline is the heavy one:
- String encryption. Every string literal becomes a call to an injected decoder. XOR mode encrypts the bytes with a key and prepends a tiny
__xhelper; the alternative builds a rotated string array with a__sdecoder. Both mean astringsdump of your script shows nothing useful. - Arithmetic encoding. Numeric literals become expressions.
5005turns into3899+1106. Pure noise, but it defeats naive constant scanning. - Control-flow flattening. Straight-line code becomes a dispatch loop over a state variable, so reading the output top-to-bottom tells you nothing about execution order.
- Dead-code injection. Plausible-looking branches that never execute, added because the cost of proving they never execute is what static analysis spends its time on.
- Identifier renaming and bracketized properties.
app.use(...)becomes_0x26f614["use"](...), and every binding gets a hex-style name.
A real before/after from the test fixtures:
// before
const app = express();
app.use(cors({ origin: ['https://your-app.example.com'], credentials: true }));
// after (XOR mode, formatted for reading)
function __x(_0xfe6fba, _0x560f77) {
var _0x8b4db9 = '';
for (var _0xf0a79b = 0; _0xf0a79b < _0xfe6fba["length"]; _0xf0a79b++)
_0x8b4db9 += String["fromCharCode"](_0xfe6fba[_0xf0a79b] ^ _0x560f77);
return _0x8b4db9;
}
const _0x26f614 = _0xad2f47();
_0x26f614["use"](_0xf34a6c({
origin: [__x([9- -68, 127-46, 142-61, /* ... */], 54-17)],
credentials: true
}));
That’s actual output, only re-wrapped so it fits. In production it’s one compact line with no comments.
The fast mode taught me what the slow mode was for
The default pipeline is expensive, and for big scripts the flatten-and-inject passes dominate the runtime. Kitsune also ships a --perf mode that swaps them for cheaper transforms: identifiers get renamed to homoglyphs (Cyrillic look-alikes, which are valid identifiers and pass every eyeball test), strings get split into concatenated chunks, and numbers get factored differently.
The interesting discovery was that this isn’t just a speed knob. Flattening is structural: it’s loud, it changes stack traces, and it makes debugging your own output genuinely painful if you need to verify behavior. Homoglyph renaming is almost invisible structurally, and for cases where your threat model is “a competitor skims the code” rather than “a determined reverser with a week”, it’s the better trade. Two modes exist because those are two different problems.
Things I got wrong along the way
- The self-defending check must survive beautification. The anti-tamper helper verifies its own source at runtime, which means it breaks if anyone reformats the output. That’s the feature working, but it also means you can’t pretty-print
main_out.jsto inspect it without breaking it. Debug accordingly. - Dynamic imports and reserved names interact. The injected helpers are generated as source strings and re-parsed into the AST before every other pass, so the rename passes have to treat them as already-handled. Getting the ordering wrong meant double-encoding the helpers on a second run.
- Output is always plain JS. Parsing accepts TS and JSX, but Babel’s generator drops types on emit. That’s fine for GAS (plain JS anyway) but worth knowing before you feed it a
.tsxfile and expect types back.
Why a fox
The name and the ASCII-fox banner are partly branding and partly a promise. 狐に化かされる means “to be bewitched by the fox”: the code looks strange, but it runs exactly the way it did before. That last part is the contract. Obfuscation that changes behavior isn’t protection, it’s a different program.
A browser lite edition runs the XOR string-encryption mode client-side on the live demo page if you want to see the transform on your own code without installing anything.
thanks for reading ~say hi →