- Home
- Image Tools
- Image Text Editor
Image Text Editor
An image text editor that works two ways: scan a picture to edit the text already in it, matching colour, size and position, or add fresh text layers of your own. It runs in your browser, so the image is never uploaded.
Choose an image and it is scanned automatically. Every line of text it finds becomes editable where it sits. Or press Add text to place your own.
Exports at 1200x630 (transparent). Source images larger than 4096px on the long side are scaled down.
Everything runs in your browser. Your image is never uploaded, logged, or stored.
How to Edit Text in an ImageWhy editing text baked into a photo needs OCR and inpainting, what that costs in megabytes, and the four approaches that actually work.How to use the Image Text Editor
Choose an image, or skip it
Pick a file or drag one in and the scan starts by itself. Without an image you get a transparent canvas, which is what you want if you are exporting lettering to use elsewhere.
Tap a line and type
Every line found becomes editable where it sits. Tap it on the picture and type straight over it, with the caret in the original's own colour and size. Until you change something, the image is untouched.
Drag a box around anything else
Handwriting, a signature, a stamp, or a script the reader has no model for: drag a box around it on the image. Whatever is inside is removed, its colour and size are measured, and a cursor appears in its place. Press Add text instead for a fresh layer of your own.
Adjust on the image
Colour, size, bold and delete sit in a small bar beside the text you selected, so you can see the result without looking away. Font, weight and everything else are in the panel, with the rarely-needed controls behind More options. Drag to move, or nudge with the arrow keys.
Export
Download as PNG to keep transparency, or JPEG for a smaller file. The export renders at the image's full resolution, not at the preview size.
What happens when you drop an image in
The scan starts on its own. Every line of text it finds becomes an editable layer sitting exactly where the original sat, already carrying the original colour, size and weight. Tap one on the picture and type. The words you replace are removed from the image itself, so a shorter replacement leaves nothing showing behind it, and anything you do not touch stays exactly as it was.
The text engine is about 6.7 MB and downloads the first time a scan runs, then stays cached. It is served from this site rather than a third party, and it runs on your device: scanning uploads nothing, the same as everything else here.
Lines, not words
Recognition works in words, and handing those straight over produces a mess. On a payment receipt a single date arrives as six separate things to edit: Jul, 20, 2026, and so on. Nobody thinks of a date that way. Words are the right unit to detect and the wrong unit to edit, so they are reassembled into lines before you see them, using the gap between them. A word space runs well under one line height. The gap to something unrelated, like a logo sitting to the left of a label, runs to several, which is what keeps the two out of each other’s layers.
What it recovers, and how closely
Measured against a receipt built with known values, then read back through the tool itself. These are results, not estimates:
- Colour: exact. All six lines came back at the source hex, #6b7280 for grey labels and #111827 for the bold value, with no drift at all. Getting there needed two corrections. Averaging every dark pixel returns a colour lightened by the antialiased rim of each glyph, so only the half furthest from the paper is measured. And deciding which pixels are the lettering by area fails on dense capitals: bold all-caps put down more ink than background, and the first version read the colour off the paper and drew near-white on white.
- Size: within 0.8px on average across sizes from 22px to 44px. Solved from the height of the glyphs rather than the height of the detected box, which carries padding that varies line by line, and both heights are read off thresholded pixels rather than one off pixels and one off font metrics.
- Position: exact, taken from the glyph bounds rather than the box around them.
- Weight: stems within 6.7% of the original’s, and total ink within 6.6%. Deliberately not the same weight NUMBER as the original. Arial at 400 is heavier than the interface stack at 400, so a fit that correctly recovered “400” drew a line noticeably thinner than the one it replaced. What the eye compares is how much ink is on the page, so the weight is chosen to match that, and the dropdown often reads one step above what the source nominally was.
- Spacing: ink width within 0.2%. Tracking is solved so the replacement occupies the original’s footprint exactly, which is what stops a longer word from pushing into its neighbour and a shorter one from leaving a gap.
- Typeface: not detected, and every scanned layer starts in the interface stack instead. This is the one attribute that cannot be measured this way. The same word sets to an ink box of 4.28 in Helvetica and 4.29 in Georgia; a serif-specific test gave serif 2.29 to 2.86 against sans 1.94 to 2.46, which overlaps. Guessing anyway is worse than not guessing, because a receipt redrawn in Georgia is obviously wrong. Screenshots are set in the platform interface font, so that is where every layer starts. The dropdown changes it in one click.
Nothing changes until you change it
A scan is a proposal, not an edit. It marks out what could be changed and leaves the picture exactly as it arrived. Only when you tap a line and alter it does that one line’s area get rebuilt, and deleting the layer puts the original pixels back. Open a receipt, look through what was found, change one figure, and every logo, stamp and signature on it is still the photograph you uploaded.
That ordering matters more than any detection rule. An earlier version rebuilt every detected area the moment a scan finished, which meant simply opening an image rewrote parts of it, including lettering nobody intended to touch. No amount of cleverness about what counts as a logo makes that acceptable; not acting until asked does.
On top of that, two kinds of area are never offered at all. Lettering on saturated brand colour is refused, because paper is not coloured, and text too faint against its background is refused, because that is what a watermark is. Fed an image that is nothing but dark lettering on a yellow badge, the tool declines all of it, says so, and leaves the badge measuring #f5b301, the exact colour it started as.
A third rule used to exist and was removed after testing against a photographed form. It refused any area whose surroundings varied a lot, on the theory that variation means a logo. A watermark printed across the whole page, plus camera noise and a shadow falling over one corner, makes every line on such a document look varied, so the rule refused most of a receipt it should have been reading. Variation now only affects the warning about how visible a patch will be, which is the thing it actually measures.
Handwriting, signatures, and scripts it cannot read
Automatic detection reads print. It does not read handwriting, and no setting changes that: the model was trained on scanned documents, and a filled-in form comes back with its printed labels found and its handwritten entries missed. Tested on a photographed tailor’s receipt, none of the four handwritten fields were recognised while eleven printed ones were.
So there is a second way in, and it works on anything at all. Drag a box around whatever you want gone and it becomes an editable area like any other: the contents are removed, the colour and size are measured from what was inside, and a cursor appears in its place. Handwriting, a signature, a stamp, a price sticker, a script with no model behind it. The box does not care what it contains.
Two honest limits. The colour is measured from everything inside the box, so drawing tightly around the writing gives a better match than sweeping in a printed rule beside it, and the colour control sits right there if the guess is off. And the replacement is type. It will land in the pen’s own colour at the pen’s own size, and it will still look like a font, because no typeface is a particular person’s hand. Reproducing that needs handwriting synthesis, which is a generative model and a different kind of tool.
What it will not do
It will not reconstruct a complicated background. Removing a word means extending the pixels beside it inwards, row by row, so it follows a gradient, a ruled line or a tinted band without leaving a seam. What it cannot do is invent detail. Words sitting over brick, foliage or a face will leave a mark, and the tool counts those areas and warns you rather than letting you find out in the export. Rebuilding texture that was never captured needs a generative model, and that is the part this deliberately does not carry.
It also will not identify the typeface. Everything else about the original is measurable from the pixels: where it sits, how big it is, what colour, roughly how heavy. The typeface is not, and the numbers behind that are in the section above. Those two limits are the honest edges of what an image text editor can do with no server behind it, and knowing where they fall is more useful than a promise that ignores them.
Your image never leaves this tab
The file you choose is decoded by your own browser, drawn to a canvas element, and exported with the browser’s own encoder. There is no upload step, no queue, and no copy on a server, because the bytes never travel anywhere they could be stored.
That matters more for photos than for most file types. Screenshots carry account names and message threads; camera files carry faces and often location metadata. Every competing image text editor examined uploads the picture to process it, and two of them say in their own terms how long they keep it.
| Tool | Uploads your photo | Cost | Keeps it for |
|---|---|---|---|
| photext.ai | Yes | Free tier, then paid | 3–30 days |
| imagetextedit.com | Yes | Credits from $1.99 | Temporarily |
| Canva / Fotor | Yes | Account required | Stored in account |
| This tool | No | Free, no account | Nothing |
Making text readable on a busy photo
The usual failure is white text over a photo that is pale in one corner and dark in another. Half the sentence disappears. Three controls under More options fix it, and the first matters most.
Outline puts a thin dark stroke around light text and keeps it legible against anything, which is why broadcast subtitles have used the same trick for decades. It is off by default, and that default is deliberate: an outline is right for a caption over a photograph and wrong for a replacement inside a screenshot, where it makes the new words look hollow and pasted on while the type around them stays solid. Shadow is softer and works better on gradients. Opacity below about 80% lets the photo through and almost always costs more readability than it buys.
Size is expressed as a percentage of image height rather than in pixels. That is deliberate: the same layer then looks identical whether the source is a 600px thumbnail or a 4000px photo, and the preview you are looking at matches the file you download.
Text on a transparent background
Skip the image entirely and you get a transparent canvas, shown as a checkerboard. Add text, export as PNG, and the result is lettering with a clear background you can drop onto anything: a video edit, a slide, a different design tool.
The format choice is the part people get wrong. PNG stores an alpha channel and keeps transparency. JPEG has no alpha at all, so exporting a transparent canvas as JPEG fills the empty area. This tool paints it white rather than letting it come out black, but the transparency is gone either way. If you need a clear background, the answer is always PNG.
Frequently asked questions
Can this edit text that is already in an image?
Yes. Choose an image and it is scanned automatically, and each line of text in it becomes an editable layer matched to the original's colour, size, weight and position. Tap one on the picture and type over it. The original is removed from the image rather than covered, so a shorter replacement leaves nothing showing behind it. That is clean on screenshots, receipts and documents, where the background behind the text is flat, and not clean on a photograph where words sit over texture, because rebuilding detail that was never captured needs a generative model.
Does scanning upload my image?
No. The text engine is about 6.7 MB and downloads to your browser the first time a scan runs, then stays cached. It is served from this site rather than a third party, and recognition runs on your own device. The image itself never leaves the tab, scanned or not.
Will it change my logo or the watermark?
No, and the main reason is ordering rather than detection. A scan only marks out what could be changed; the picture is untouched until you edit a particular line, and only that line's area is rebuilt. Deleting a layer restores the original pixels. On top of that, two kinds of area are never offered at all: lettering on saturated brand colour, because paper is not coloured, and text too faint against its background, which is what a watermark is. Given an image that is nothing but dark lettering on a yellow badge, the tool refuses all of it and leaves the badge at the exact colour it started as.
Can it edit handwriting?
It can replace handwriting but not read or imitate it. Automatic detection is trained on printed documents, so a filled-in form comes back with its printed labels found and its handwritten entries missed. Tested on a photographed tailor's receipt, none of the four handwritten fields were recognised. For those, drag a box around the writing: whatever is inside is removed, the colour and size are measured from it, and you type in its place. The result lands in the pen's own colour at roughly its own size, and it will still look like type, because no font is a particular person's handwriting. Reproducing that needs handwriting synthesis, which is a generative model and a different kind of tool.
Does it work with languages other than English?
Partly. The bundled recognition model is English, so other scripts are usually not detected automatically. What did change is that non-Latin text is no longer thrown away when it is detected: the filter used to require a Latin letter or digit, which silently refused Urdu, Arabic, Chinese and Cyrillic as though they were not text at all. For any script the reader misses, the box tool works regardless of language, because it measures pixels rather than reading them.
Does it detect the original font?
It matches colour exactly, size to within about a pixel, and position from the glyph bounds. Weight and spacing are matched by appearance rather than by number: comparing replaced lines against the pixels they replaced, stems land within 6.7% of the original's thickness and the line occupies its width to within 0.2%. The weight shown in the panel is often a step above what the source nominally was, because the interface stack is lighter than the typefaces most screenshots use. It does not identify the typeface, and starts every scanned layer in the interface stack instead, which is what screenshots are set in. That limit is measured rather than assumed: the same word produces an ink box of 4.28 in Helvetica and 4.29 in Georgia, which is indistinguishable, and a serif-specific test overlapped as well. Identifying a typeface needs a model trained on glyph shapes, which would weigh more than everything else here combined. The dropdown changes it in one click.
Is my image uploaded anywhere?
No. The file is decoded by your browser, drawn to a canvas, and exported by the browser's own encoder. There is no upload, no server copy, and nothing retained after you close the tab. That is worth knowing for screenshots and camera photos, which often carry more personal information than people expect. An image text editor that runs entirely in the tab removes the question rather than answering it in a privacy policy.
How do I get text with a transparent background?
Add text without choosing an image, then export as PNG. The canvas starts transparent and PNG stores an alpha channel, so the download is lettering on a clear background. JPEG cannot do this, because it has no alpha channel, so any transparent area gets filled in.
Why does my text disappear against part of the photo?
Because the background changes brightness underneath it. Open More options and increase the outline width, which puts a dark edge around light text and keeps it readable over anything, or add a shadow for a softer effect on gradients. Outline is off by default on purpose: it is right for a caption over a photograph and wrong inside a screenshot, where it makes replaced words look hollow next to the solid type around them.
Does the download match what I see in the preview?
Yes. Layer positions and sizes are stored as proportions of the canvas rather than fixed pixel values, so the same layer renders identically at preview size and at the image's full resolution. The export is generated at the source image's dimensions, not at the size shown on screen.
Which fonts are available, and can I upload my own?
Six families are offered: interface, sans, serif, mono, condensed and display, all built from typefaces already present on your system. Interface is the default because it resolves to whatever your platform uses for its own menus and buttons, which is what the screenshots people bring here are set in. Custom font upload is not supported, deliberately: loading webfonts would add hundreds of kilobytes to a tool that currently ships a few, and the page would get slower for everyone to serve a minority of uses.
Is there a size limit on the image?
Images larger than 4096 pixels on the longest side are scaled down to that, and the page tells you when it happens. This keeps the canvas within reach of a phone's memory. A modern 50-megapixel camera file would otherwise allocate hundreds of megabytes and stall the browser. Exports use the scaled dimensions, which are still far larger than any screen use requires.